Challenges and Future Directions#
Recent LM-based VAD methods have expanded the scope of anomaly detection by introducing semantic reasoning, open-vocabulary recognition, and natural-language explanation. At the same time, a growing line of work has begun to examine the limitations and vulnerabilities introduced by this shift. In particular, recent analyses argue that the increasing reliance on multi-scene formulations, weak supervision, and pretrained LLM/MLLM priors may move VAD away from its original goal of modeling scene-specific deviations from normality [161]. These concerns motivate several open challenges for future LM-based VAD research.
Scene-specific normality and the limits of category-based anomaly detection. A central challenge for LM-based VAD is that the semantic knowledge provided by LLMs and MLLMs can encourage a category-based view of anomalies. Many recent methods are effective at recognizing common abnormal event categories, such as fighting, burglary, fire, accidents, or falling. However, anomaly detection is not equivalent to recognizing a predefined set of abnormal actions. The same activity may be normal in one scene and anomalous in another, or even normal in one region of a scene and anomalous in a different region of the same scene. For example, fighting inside a boxing ring is expected, whereas fighting among the audience is anomalous. This distinction highlights the need to model normality as a scene-specific and context-dependent property rather than as a fixed semantic label. Future LM-based VAD should therefore move beyond detecting familiar anomaly categories and instead reason about whether an observed event violates the normal activity patterns of the target environment.
Spatial grounding, localization, and context-dependent reasoning. Many LM-based VAD methods, especially weakly supervised and training-free approaches, operate primarily at the video or frame level. While this formulation is useful for reporting whether an anomaly occurs, it often fails to identify where the anomaly occurs and which objects, people, or interactions are responsible. This limitation is especially important for anomalies whose abnormality depends on spatial context, such as illegal parking, jaywalking, entering a restricted area, abandoning an object, or interacting with an object in an unusual location. Without spatial grounding, a system may detect an anomalous frame but still fail to provide actionable information to a human operator. Moreover, reliable explanation generally requires localization: a model cannot faithfully explain an anomaly without identifying the visual evidence that supports its decision. Future work should therefore place greater emphasis on object-level, region-level, and track-level anomaly localization, together with evaluation protocols that measure whether the detected anomaly is spatially aligned with the true abnormal event.
Semantic bias, hallucination, and evaluation transparency. The use of large pretrained language and vision-language models introduces both opportunities and risks. These models provide useful semantic priors, but their prior knowledge may dominate the anomaly decision. A model may classify an event as anomalous because the action is generally associated with abnormality in its pretraining distribution, even when the event is normal in the current scene. Conversely, subtle scene-specific anomalies may be missed if they do not correspond to familiar anomaly categories. MLLMs may also generate plausible but incorrect explanations, including hallucinated objects, actions, or causal relations. This issue is further complicated by the opacity of pretrained models, since their training data are often undisclosed. As a result, it is difficult to determine whether benchmark videos, anomaly classes, or visually similar events have been indirectly observed during pretraining. Future evaluations should therefore assess not only detection accuracy, but also explanation faithfulness, calibration, grounding, and robustness to unseen anomaly types.
Efficient, explainable, and adaptive LM-based VAD systems. Future LM-based VAD systems should combine the semantic flexibility of language models with explicit models of normality learned from the target scene. Rather than using LMs only as generic anomaly classifiers, future systems can use them as reasoning modules that inspect localized evidence, compare observations against scene-specific normality rules, retrieve similar normal events, and generate grounded explanations. Such systems should also be efficient and extendable, since real deployments may involve many cameras and continuously evolving normal activity patterns. This motivates hybrid designs in which lightweight perception modules perform detection, tracking, and feature extraction, while smaller or task-specialized language models support rule induction, contextual reasoning, explanation, and updating. More broadly, the rise of agentic and dynamic systems points toward VAD frameworks that are not static detectors, but adaptive reasoning pipelines capable of refining their decisions and explaining anomalies in relation to the evolving normality of each scene.
Efficiency, scalability, and real-time deployment. Another major challenge is the computational cost of LM-based VAD. Many recent methods rely on large VLMs, MLLMs, captioning models, or repeated LLM queries, which can be expensive when applied to long videos or multi-camera surveillance systems. This limits their practicality in real-time settings, where anomaly scores, localization outputs, and explanations may need to be generated with low latency. The challenge becomes even more significant when models require dense frame sampling, long-context reasoning, or multiple rounds of prompting. Future work should therefore explore efficient architectures that combine lightweight visual backbones, object detectors, trackers, retrieval modules, and smaller task-specialized language models. A promising direction is to use low-cost perception modules for continuous monitoring and invoke larger reasoning models only when uncertain or suspicious events are detected. Such hierarchical designs can preserve the semantic advantages of LMs while improving scalability, latency, and deployability.
Toward agentic and dynamic VAD systems. The recent rise of agentic AI has introduced new ways to design systems that can plan, retrieve information, use tools, verify intermediate outputs, and refine decisions through multi-step reasoning [151], [152], [162]-[164]. This trend suggests a promising direction for LM-based VAD. Instead of treating the language model as a passive classifier, caption generator, or anomaly scorer, future systems may use LMs as active reasoning agents that decompose anomaly detection into multiple steps. For example, an agentic VAD system could observe the scene, retrieve normal examples from memory, identify relevant objects and interactions, check location-specific rules, compare current activity with historical patterns, and produce grounded explanations. Such systems could also support dynamic adaptation, where normality models are updated as new normal activities appear or as the environment changes over time. In this view, VAD systems would no longer operate as static detectors, but as adaptive reasoning pipelines capable of refining their assumptions and explaining anomalies with respect to evolving scene-specific normality.
Ethical, privacy, and societal considerations. The deployment of LM-based VAD raises important ethical and societal concerns. Since VAD is often associated with surveillance, models that generate natural-language descriptions of people, actions, and interactions may increase privacy risks and create new opportunities for misuse. Language-based explanations can also make systems appear more reliable than they actually are, especially when generated explanations are plausible but not visually grounded. Bias is another concern: models may learn or inherit assumptions about what constitutes suspicious behavior, leading to unfair or context-insensitive judgments. Beyond privacy and bias, robustness is also critical. Recent studies have shown that VAD systems can be vulnerable to adversarial perturbations, distribution shifts, and intentionally manipulated inputs [165], [166]. These vulnerabilities are especially concerning in safety-critical or security-sensitive environments, where an attacker may attempt to hide anomalous behavior or trigger false alarms. Future LM-based VAD research should therefore incorporate robustness evaluation, adversarial testing, privacy-preserving designs, human-in-the-loop verification, and transparent reporting of failure cases.