Training-Free From survey PDF — verify links
LAVAD
LAVAD — Caption-Based Temporal Reasoning
CVPR · 2024Human-reviewed summary status
Survey summary
LAVAD is a training-free method that turns VAD into language-based reasoning: a VLM captioning model describes each frame, an LLM aggregates captions over temporal windows to estimate anomaly scores, and cross-modal similarity cleans noisy captions and refines the scores.
Taxonomy placement
Training-Free
Training-FreeCaption-Based Temporal Reasoning
Method characteristics
Training-FreeCaption-basedLLM reasoningCross-modal refinement