Survey Scope and Literature Collection#
This survey provides a focused and structured review of language-model-based video anomaly detection. We distinguish between two literature collections used for different purposes in the paper: (i) a fixed-venue corpus used for the publication-trend analysis in Figure 1 and Table 1, and (ii) a broader literature corpus used for the technical survey and taxonomy.
For the publication-trend analysis, we exhaustively reviewed VAD papers published in the major computer vision and machine learning conferences considered in this work: CVPR, ICCV, ECCV, WACV, NeurIPS, and ICML. All papers whose primary contribution concerns video anomaly detection were inspected, and LM-based papers were identified according to the operational criteria in Table 7. The cutoff for this fixed-venue analysis is the end of the 2025 conference cycle. This corpus is used only for the venue-level publication statistics reported in Figure 1 and Table 1.
For the broader technical survey, we searched major scholarly search engines, publisher databases, and conference repositories, including Google Scholar, arXiv, IEEE Xplore, ACM Digital Library, and the CVF Open Access repository. Searches used combinations of terms including *video anomaly detection*, *video anomaly understanding*, *language model video anomaly detection*, *vision-language video anomaly detection*, *multimodal large language model anomaly detection*, *CLIP video anomaly detection*, *training-free video anomaly detection*, and *open-vocabulary video anomaly detection*, together with closely related variants. Reference lists of relevant papers and recent VAD surveys were additionally inspected to identify potentially missed works. The broader survey corpus includes literature available up to May 2026.
A work was considered within the primary scope of the survey when anomaly-oriented video analysis constitutes a central contribution and language models, vision-language models, multimodal large language models, language-aligned representations, or language-derived supervision play a substantive methodological role. We consider not only conventional anomaly detection and localization, but also emerging anomaly-oriented tasks such as anomaly understanding, explanation, retrieval, question answering, and grounding when these are explicitly formulated around abnormal-event analysis. In contrast, image-only anomaly detection, generic video understanding or action recognition, and surveillance analysis without an explicit anomaly-oriented objective are excluded from the primary survey taxonomy.
Both peer-reviewed publications and technically substantive preprints are included in the broader survey corpus because a substantial portion of recent LM-based VAD research appears first on arXiv before formal publication. When a peer-reviewed version is available, it is preferred over the corresponding preprint, and duplicate preprint and published versions are not treated as separate works. Candidate papers were first screened based on title and abstract and were subsequently inspected at the full-paper level to confirm their relevance and methodological role.
To make the survey boundaries and classification procedure explicit, Table 7 summarizes the operational criteria used to determine whether a work belongs to the primary LM-based VAD corpus and how representative methodological paradigms are assigned. These categories are not necessarily mutually exclusive; rather, they identify the principal methodological role used to organize the literature. Future additions to the dynamic survey are screened according to the same scope and classification principles before being incorporated into maintained releases.
TABLE 7. Operational inclusion and classification criteria used in constructing the LM-based VAD survey corpus.
| Criterion | Inclusion / positive rule | Exclusion boundary |
|---|---|---|
| LM-based VAD paper | Included when video anomaly detection, localization, understanding, explanation, retrieval, or a closely related anomaly-oriented video task is a primary contribution, and an LLM, VLM, MLLM, language-aligned representation, or language-derived supervision plays a substantive methodological role. | Image-only anomaly detection, generic video understanding or action recognition, and surveillance analysis without an explicit anomaly-oriented objective are excluded from the primary taxonomy. |
| Publication-trend corpus | Includes all VAD papers identified in CVPR, ICCV, ECCV, WACV, NeurIPS, and ICML through the end of the 2025 conference cycle; LM-based papers are labeled using the criteria in this table. | Papers outside these venues are not included in the statistics of Figure 1 and Table 1, even if they are included elsewhere in the broader survey. |
| Publication status | Both peer-reviewed papers and technically substantive preprints are considered in the broader survey corpus. When a peer-reviewed version is available, it is preferred over the corresponding preprint. | Duplicate preprint and published versions are not treated as separate works. |
| Language-model involvement | Classified as LM-based when language representations, textual prompts, language-model reasoning, textual supervision, or multimodal language models directly contribute to anomaly detection, localization, or understanding. | A paper is not classified as LM-based merely because an LLM is used for auxiliary writing, metadata generation, or other components unrelated to the anomaly-analysis mechanism. |
| Training-free | Classified as training-free when pretrained model parameters remain fixed with respect to the VAD task and anomaly inference is performed through prompting, similarity estimation, retrieval, captioning, or frozen-model reasoning. | Methods that perform VAD-specific parameter optimization or fine-tuning are not classified as training-free. |
| Instruction-tuned | Classified as instruction-tuned when a pretrained LM, VLM, or MLLM is adapted using anomaly-related instruction, question-answer, explanation, grounding, or similar task-specific supervision. | Prompting a frozen model without parameter adaptation does not satisfy this criterion. |
| Open-world / open-vocabulary | Classified in this category when the method explicitly supports previously unseen anomaly categories, expandable textual label spaces, or user-defined anomaly concepts at inference time. | Standard closed-set anomaly detection using a fixed training-category space is not classified as open-world or open-vocabulary. |
| Anomaly understanding / semantic output | Included when semantic outputs such as explanations, question-answer responses, anomaly categories, retrieval results, captions, or grounding predictions are explicitly evaluated as part of an anomaly-oriented task. | Generic video captioning, retrieval, or question answering without an anomaly-centered objective is excluded from the primary corpus. |