← Back to literature explorer
Training-Free From survey PDF — verify links

LAVAD

LAVAD — Caption-Based Temporal Reasoning

CVPR · 2024

Survey summary

LAVAD is a training-free method that turns VAD into language-based reasoning: a VLM captioning model describes each frame, an LLM aggregates captions over temporal windows to estimate anomaly scores, and cross-modal similarity cleans noisy captions and refines the scores.

Training-Free

Training-FreeCaption-Based Temporal Reasoning
Training-FreeCaption-basedLLM reasoningCross-modal refinement