Challenges and Future Directions#
Recent LM-based VAD methods have expanded the scope of anomaly detection by introducing semantic reasoning, open-vocabulary recognition, and natural-language explanation. At the same time, a growing line of work has begun to examine the limitations and vulnerabilities introduced by this shift. In particular, recent analyses argue that the increasing reliance on multi-scene formulations, weak supervision, and pretrained LLM/MLLM priors may move VAD away from its original goal of modeling scene-specific deviations from normality [169]. These concerns motivate several open challenges for future LM-based VAD research.
Scene-specific normality and the limits of category-based anomaly detection. A central challenge for LM-based VAD is that the semantic knowledge provided by LLMs and MLLMs can encourage a category-based view of anomalies. Many recent methods are effective at recognizing common abnormal event categories, such as fighting, burglary, fire, accidents, or falling [45], [101], [127], [128], [130], [138], [148]. However, anomaly detection is not equivalent to recognizing a predefined set of abnormal actions. The same activity may be normal in one scene and anomalous in another, or even normal in one region of a scene and anomalous in a different region of the same scene [27], [51]. For example, fighting inside a boxing ring is expected, whereas fighting among the audience is anomalous. This distinction highlights the need to model normality as a scene-specific and context-dependent property rather than as a fixed semantic label. Future LM-based VAD should therefore move beyond detecting familiar anomaly categories and instead reason about whether an observed event violates the normal activity patterns of the target environment.
Spatial grounding, localization, and context-dependent reasoning. Many LM-based VAD methods, especially weakly supervised and training-free approaches, operate primarily at the video or frame level [101], [110], [126], [138], [139]. While this formulation is useful for reporting whether an anomaly occurs, it often fails to identify where the anomaly occurs and which objects, people, or interactions are responsible. This limitation is especially important for anomalies whose abnormality depends on spatial context, such as illegal parking, jaywalking, entering a restricted area, abandoning an object, or interacting with an object in an unusual location [50], [51]. Without spatial grounding, a system may detect an anomalous frame but still fail to provide actionable information to a human operator. Moreover, reliable explanation generally requires localization: a model cannot faithfully explain an anomaly without identifying the visual evidence that supports its decision. Future work should therefore place greater emphasis on object-level, region-level, and track-level anomaly localization, together with evaluation protocols that measure whether the detected anomaly is spatially aligned with the true abnormal event.
Semantic bias, hallucination, and evaluation transparency. The use of large pretrained language and vision-language models introduces both opportunities and risks. These models provide useful semantic priors, but their prior knowledge may dominate the anomaly decision [127], [128], [169]. A model may classify an event as anomalous because the action is generally associated with abnormality in its pretraining distribution, even when the event is normal in the current scene. Conversely, subtle scene-specific anomalies may be missed if they do not correspond to familiar anomaly categories. This semantic-prior bias can therefore manifest in both directions: false positives may arise when generally unusual but scene-appropriate activities are treated as anomalous, while false negatives may occur when context-dependent anomalies do not match familiar anomaly concepts. Scene and background correlations may further bias decisions toward camera, location, or contextual cues rather than the anomalous event itself.
MLLMs may also generate plausible but incorrect explanations, including hallucinated objects, actions, or causal relations [170], [171]. Prompt sensitivity represents another source of instability, since changes in prompt wording, contextual information, or candidate anomaly descriptions may alter the resulting anomaly judgment. Similarly, limited temporal reasoning or sparse frame sampling can cause models to miss anomalies that depend on motion, event ordering, or longer-term temporal context. These issues are particularly important for LM-based VAD because the anomaly decision may depend jointly on visual evidence, temporal context, and language-based reasoning.
This issue is further complicated by the opacity of pretrained models, since their training data are often undisclosed [172]. As a result, it is difficult to determine whether benchmark videos, anomaly classes, or visually similar events have been indirectly observed during pretraining [173], [174]. Such potential benchmark contamination complicates the interpretation of zero-shot and open-world performance, particularly when the extent of overlap between pretraining data and evaluation benchmarks cannot be established. More broadly, contextual and cultural assumptions inherited from large-scale pretraining may influence what a model considers suspicious or abnormal, which is problematic when acceptable behavior varies across scenes, environments, or deployment contexts.
Another open challenge is the lack of a unified evaluation benchmark for language-based VAD outputs. Existing anomaly-understanding datasets differ substantially in task formulation, annotation style, output format, and evaluation protocol, spanning question answering, explanation generation, temporal grounding, retrieval, and captioning. As a result, semantic performance reported across these benchmarks is generally not directly comparable. Unlike conventional frame-level VAD, where metrics such as AUC and AP are widely established, there is currently no broadly adopted protocol for jointly evaluating the correctness, grounding, faithfulness, and usefulness of language-based anomaly explanations or reasoning. Developing standardized benchmarks and evaluation procedures, potentially combining human assessment with reproducible LLM/MLLM-based judging, is therefore an important direction for future work.
Future evaluations should therefore assess not only detection accuracy, but also explanation faithfulness, calibration, grounding, and robustness to unseen anomaly types. In addition, they should examine prompt robustness, scene dependence, temporal failure cases, false-positive and false-negative behavior, and performance under distribution shifts or intentionally manipulated inputs.
Efficient, explainable, and adaptive LM-based VAD systems. Future LM-based VAD systems should combine the semantic flexibility of language models with explicit models of normality learned from the target scene. Rather than using LMs only as generic anomaly classifiers, future systems can use them as reasoning modules that inspect localized evidence, compare observations against scene-specific normality rules, retrieve similar normal events, and generate grounded explanations. Such systems should also be efficient and extendable, since real deployments may involve many cameras and continuously evolving normal activity patterns. This motivates hybrid designs in which lightweight perception modules perform detection, tracking, and feature extraction, while smaller or task-specialized language models support rule induction, contextual reasoning, explanation, and updating [79], [81]. More broadly, the rise of agentic and dynamic systems points toward VAD frameworks that are not static detectors, but adaptive reasoning pipelines capable of refining their decisions and explaining anomalies in relation to the evolving normality of each scene.
Efficiency, scalability, and real-time deployment. Another major challenge is the computational cost of LM-based VAD. Many recent methods rely on large VLMs, MLLMs, captioning models, or repeated LLM queries, which can be expensive when applied to long videos or multi-camera surveillance systems [45], [126], [130], [135], [139], [141]. This limits their practicality in real-time settings, where anomaly scores, localization outputs, and explanations may need to be generated with low latency [82], [132], [134]. The challenge becomes even more significant when models require dense frame sampling, long-context reasoning, or multiple rounds of prompting. Future work should therefore explore efficient architectures that combine lightweight visual backbones, object detectors, trackers, retrieval modules, and smaller task-specialized language models. A promising direction is to use low-cost perception modules for continuous monitoring and invoke larger reasoning models only when uncertain or suspicious events are detected. Such hierarchical designs can preserve the semantic advantages of LMs while improving scalability, latency, and deployability. Reproducibility is further complicated by the heterogeneous foundation models used across LM-based VAD systems. Some approaches rely on proprietary or rapidly evolving models, while others use open checkpoints whose versions and inference configurations may differ across studies. Important implementation details such as prompt templates, decoding parameters, number of model calls, hardware, runtime, latency, and API cost are also reported inconsistently. For this reason, direct computational comparison across the current literature is difficult. To provide a more transparent overview of model dependence, the [Reproducibility Appendix](/survey/appendix/) summarizes the foundation models used by the surveyed LM-based VAD methods.
Toward agentic and dynamic VAD systems. The recent rise of agentic AI has introduced new ways to design systems that can plan, retrieve information, use tools, verify intermediate outputs, and refine decisions through multi-step reasoning [159], [160], [175]-[177]. This trend suggests a promising direction for LM-based VAD. Instead of treating the language model as a passive classifier, caption generator, or anomaly scorer, future systems may use LMs as active reasoning agents that decompose anomaly detection into multiple steps. For example, an agentic VAD system could observe the scene, retrieve normal examples from memory, identify relevant objects and interactions, check location-specific rules, compare current activity with historical patterns, and produce grounded explanations. Such systems could also support dynamic adaptation, where normality models are updated as new normal activities appear or as the environment changes over time. In this view, VAD systems would no longer operate as static detectors, but as adaptive reasoning pipelines capable of refining their assumptions and explaining anomalies with respect to evolving scene-specific normality.
Ethical, privacy, and societal considerations. The deployment of LM-based VAD raises important ethical and societal concerns. Since VAD is often associated with surveillance, models that generate natural-language descriptions of people, actions, and interactions may increase privacy risks and create new opportunities for misuse [34], [178]. Language-based explanations can also make systems appear more reliable than they actually are, especially when generated explanations are plausible but not visually grounded. Bias is another concern: models may learn or inherit assumptions about what constitutes suspicious behavior, leading to unfair or context-insensitive judgments [179]. Beyond privacy and bias, robustness is also critical. Recent studies have shown that VAD systems can be vulnerable to adversarial perturbations, distribution shifts, and intentionally manipulated inputs [180], [181]. These vulnerabilities are especially concerning in safety-critical or security-sensitive environments, where an attacker may attempt to hide anomalous behavior or trigger false alarms. Future LM-based VAD research should therefore incorporate robustness evaluation, adversarial testing, privacy-preserving designs, human-in-the-loop verification, and transparent reporting of failure cases.
Real-world deployment also requires attention to operational factors that are not captured by benchmark accuracy alone. In large-scale surveillance systems, even relatively low false-positive rates can generate substantial numbers of alerts, making false-alarm management and human review important components of practical system design. Deployment policies should also consider how video data, generated descriptions, and model outputs are stored or retained, particularly because language-based representations may expose sensitive information that is less apparent in raw anomaly scores. In addition, bandwidth, edge-computing constraints, and reliance on remote or proprietary foundation-model services can affect where and how LM-based VAD can be deployed. These considerations motivate evaluation frameworks that account not only for detection performance, but also for operational reliability, human oversight, privacy, and deployment-specific resource constraints.