Survey Versionv1.0Last UpdatedJuly 9, 2026Last Changed Inv1.0

Conclusion#

Language models are rapidly changing the landscape of video anomaly detection by extending the field beyond visual pattern recognition toward semantic reasoning, contextual interpretation, open-vocabulary recognition, and natural-language explanation. In this survey, we reviewed this emerging research direction through a unified taxonomy of language-model-based VAD methods, covering unsupervised and semi-supervised, weakly supervised, training-free, instruction-tuned, and open-world/open-vocabulary paradigms. We also positioned these methods relative to key non-language-model foundations, reviewed representative datasets and evaluation protocols, and discussed how recent benchmarks increasingly shift the focus from anomaly detection alone toward anomaly understanding.

Across the surveyed literature, a clear trend emerges: language models provide powerful semantic priors that can help VAD systems reason about objects, actions, interactions, scene context, and unseen anomaly concepts. However, this shift also introduces new challenges. LM-based VAD systems must avoid reducing anomaly detection to category recognition, since abnormality is often scene-specific and context-dependent. They must also provide reliable spatial and temporal grounding, reduce hallucination and semantic bias, remain computationally efficient, and support transparent evaluation beyond conventional detection metrics. These challenges are especially important as VAD systems move toward real-world deployment, where explanations must be faithful, decisions must be robust, and normality may evolve over time.

Finally, because LM-based VAD is developing rapidly, we frame this work not only as a static review but also as a dynamic and versioned scholarly resource. The maintained survey will support the addition of newly published methods, datasets, benchmarks, and evaluation protocols through localized revisions, while preserving the overall taxonomy, terminology, and organizational structure of the peer-reviewed article. Each release will be accompanied by transparent version records and archived through the public website, allowing readers to identify newly added papers, revised sections, updated tables, and other substantive changes over time. This maintenance process is intended to reduce the obsolescence that commonly affects surveys in fast-moving research areas and to provide a stable reference point for the LM-based VAD community.

Acknowledgment#

The authors used ChatGPT, powered by OpenAI's GPT-5.5 model [167], to assist with the visual editing and polishing of Figs. 3-7. The figure concepts, scientific content, and final revisions were developed, reviewed, and approved by the authors.