Survey Versionv1.0Last UpdatedSeptember 21, 2026Last Changed Inv1.0

Datasets#

Table 3 summarizes representative VAD datasets and compares them according to their annotation granularity and semantic richness. While traditional VAD benchmarks primarily provide temporal and spatial anomaly annotations, recent datasets increasingly incorporate higher-level semantic supervision. To capture this distinction, we include a semantic annotation category that describes the richest form of semantic information available in each dataset.

Semantic Annotation refers to the highest level of semantic supervision provided by a dataset, ranging from anomaly category labels to question-answer pairs, grounding annotations, retrieval pairs, captioning supervision, and textual explanations.

Early VAD research was largely driven by scene-specific datasets designed to evaluate statistical anomaly detection methods. Representative benchmarks include Subway Entrance and Subway Exit [7], UMN [9], UCSD Ped1 and UCSD Ped2 [47], CUHK Avenue [20], ShanghaiTech [48], NWPU Campus [49], and Street Scene [50]. These datasets primarily provide frame-level anomaly labels and, in some cases, spatial annotations through bounding boxes. Their focus is on detecting deviations from learned normal patterns within relatively constrained environments. Among them, Street Scene [50] provides track-level temporal annotations that support object-centric anomaly localization and interaction analysis. ComplexVAD [51], a large-scale streetscape benchmark containing complex interaction anomalies with both track-level temporal annotations and bounding-box spatial annotations, enabling fine-grained object-centric anomaly localization and future research on anomaly explanation and understanding. Despite their importance for evaluating anomaly detection performance, these datasets generally lack semantic annotations and therefore provide limited support for language-based reasoning, explanation generation, or anomaly understanding.

To improve scalability and anomaly diversity, later benchmarks expanded beyond single-scene settings and introduced category-level semantic supervision. Representative examples include UCF-Crime [19], UCF-Crime Extension [52], XD-Violence [53], TAD [54], BOSS [55], CamNuvem [56], Ubnormal [57], and the dataset introduced by Zhu et al. [58]. Compared with earlier benchmarks, these datasets contain substantially larger numbers of videos and anomaly categories, enabling the development of weakly supervised and fully supervised VAD methods. However, semantic supervision remains relatively coarse, typically consisting of anomaly category labels rather than detailed descriptions, explanations, or reasoning-oriented annotations.

More recently, several benchmarks have been proposed to evaluate anomaly understanding rather than anomaly detection alone. Examples include CUVA [59], ECVA [60], VANE-Bench [61], VAGU [62], FineVAU [63], HVAU-70K [45], and UCA [64]. These datasets introduce richer semantic annotations, including question-answer pairs, conversational interactions, grounding annotations, and fine-grained reasoning tasks. Similarly, the retrieval benchmark proposed by Yang et al. [65] and the captioning benchmark introduced by Bao et al. [66] extend evaluation toward video-text retrieval and anomaly description generation. Such benchmarks enable the assessment of higher-level capabilities including anomaly explanation, contextual reasoning, temporal grounding, and language-guided understanding, making them particularly relevant for modern vision-language models and MLLMs.

Despite their importance for benchmarking, existing VAD datasets exhibit several limitations that should be considered when interpreting reported performance. Many conventional benchmarks contain data from a limited number of cameras, geographic locations, or environmental conditions, which can encourage scene-specific or background-dependent representations and limit conclusions about cross-scene generalization. Anomaly definitions may also be dataset-dependent and sometimes ambiguous, since the same behavior can be normal or abnormal depending on location, context, or local activity patterns. In addition, VAD datasets are typically highly imbalanced, with anomalous events occupying only a small fraction of the available video, while multi-scene datasets may still provide limited diversity relative to real deployment environments. Some recent benchmarks increase anomaly diversity through curated or synthetically generated content; for example, VANE-Bench includes both real-world surveillance videos and synthetically generated anomalous videos [61]. While such data can broaden anomaly coverage, synthetic or curated events may not fully capture the visual, temporal, and contextual variability of naturally occurring real-world anomalies. Closed-set category annotations can further encourage category recognition rather than scene-specific anomaly reasoning. These factors should therefore be considered when transferring benchmark results to operational settings, and future datasets would benefit from greater camera, scene, geographic, and contextual diversity together with richer annotations of ambiguous and location-dependent abnormality.

TABLE 3. Comparison of representative VAD datasets in terms of annotation granularity and semantic richness. Datasets are grouped according to their dominant annotation and semantic characteristics.

DatasetYearDomainTotal FramesAnomaly CategoriesTemporal AnnotationSpatial AnnotationSemantic Annotation
Subway Entrance [7]2008Streetscape86,5355Frame–None
Subway Exit [7]2008Streetscape38,9403Frame–None
UMN [9]2009Crowd Behavior3,8551Frame–None
UCSD Ped1 [47]2013Streetscape14,0005FrameBounding BoxNone
UCSD Ped2 [47]2013Streetscape4,5605FrameBounding BoxNone
CUHK Avenue [20]2013Streetscape30,6525FrameBounding BoxNone
NWPU Campus [49]2023Streetscape1,466,07328Frame–None
ShanghaiTech [48]2017Streetscape317,39813FrameBounding BoxNone
Street Scene [50]2020Streetscape203,25717TrackBounding BoxNone
ComplexVAD [51]2025Streetscape3,681,43840TrackBounding BoxNone
UCF-Crime [19]2018Crime13,741,39313Video–Category Labels
UCF-Crime Extension [52]2021Crime14,475,79315Video–Category Labels
XD-Violence [53]2020Violence114,0966Video–Category Labels
TAD [54]2024Traffic721,2804FrameBounding BoxCategory Labels
BOSS [55]2017Multiple48,62411Video–Category Labels
CamNuvem [56]2022Robbery6,151,7881Video–Category Labels
Ubnormal [57]2022Multiple236,90222FramePixelCategory Labels
SENSE-VAD [67]2026Autonomous Driving540,88815FrameBounding BoxCategory Labels
MSAD [58]2024Multiple447,23655Frame–Category Labels
CUVA [59]2024Multiple3,345,09711Time Duration–Video QA
ECVA [60]2024Multiple19,042,56021Time Duration–Video QA
VANE-Bench [61]2025Multiple951,48219Video–Conversational QA
VAGU [62]2025Multiple20,400,00021Period–Grounding + QA
FineVAU [63]2026Surveillance–13––Fine-grained QA
SVTA [65]2025Multiple1,360,00068Video-level–Retrieval text pairs
CVACBench [66]2025Surveillance60,16013Frame–Captioning
HVAU-70K [45]2025Multiple13,855,48915Frame–Video QA
UCA [64]2024Crime11,817,59713Frame–Video QA