FutureFrame
FutureFrame — Prediction-Based Normality
Classical non-LM baseline. Detects anomalies via large future-frame prediction errors on events that cannot be forecast from normal training video.
Search representative records by title, method, dataset, model, task, author, or venue.
| Method | Year | Paradigm | Model | Tasks | Code |
|---|---|---|---|---|---|
| FutureFrameFutureFrame — Prediction-Based Normality | 2018 | Unsupervised & Semi-Supervised | CNN | Detection | — |
| MemAEMemAE — Reconstruction-Based Normality | 2019 | Unsupervised & Semi-Supervised | Autoencoder | Detection | — |
| HyCoVADHyCoVAD — Detect-then-Validate | 2025 | Unsupervised & Semi-Supervised | SSLLLM | Detection, Explanation | — |
| MLLM-EVADMLLM-EVAD — Semantic Normality Exemplars | 2025 | Unsupervised & Semi-Supervised | MLLM | Detection, Explanation | — |
| SlowFastVADSlowFastVAD — RAG-Enhanced Reasoning | 2025 | Unsupervised & Semi-Supervised | VLMLLM | Detection, Explanation | — |
| SFN-VADSFN-VAD — Prediction + Semantic Consistency | 2025 | Unsupervised & Semi-Supervised | VLM | Detection | — |
| VLAVADVLAVAD — Semantic Normality Modeling | 2024 | Unsupervised & Semi-Supervised | VLM | Detection | — |
| Follow the Rules (AnomalyRuler)Follow the Rules (AnomalyRuler) — Rule Induction | 2024 | Unsupervised & Semi-Supervised | VLMLLM | Detection | — |
| Kim et al. (CLIP normality)Kim et al. (CLIP normality) — Text-Conditional Similarity | 2023 | Unsupervised & Semi-Supervised | CLIP | Detection | — |
| DeepMILDeepMIL — Multiple Instance Learning | 2018 | Weakly Supervised | CNN | Detection | — |
| MSLMSL — Self-Training | 2022 | Weakly Supervised | Transformer | Detection | — |
| MISTMIST — Self-Training | 2021 | Weakly Supervised | CNN | Detection | — |
| GCNGCN — Graph Reasoning | 2019 | Weakly Supervised | GCN | Detection | — |
| VadCLIPVadCLIP — Vision-Language Adaptation | 2024 | Weakly Supervised | CLIP | Detection, Classification | — |
| VadCLIP++VadCLIP++ — Vision-Language Adaptation | 2025 | Weakly Supervised | CLIP | Detection, Classification | — |
| CLIP-TSACLIP-TSA — Vision-Language Adaptation | 2024 | Weakly Supervised | CLIP | Detection | — |
| ReFLIP-VADReFLIP-VAD — Vision-Language Adaptation | 2024 | Weakly Supervised | VLM | Detection | — |
| PEMILPEMIL — Prompt-Enhanced MIL | 2024 | Weakly Supervised | VLM | Detection, Classification | — |
| TEVADTEVAD — Caption-Guided Semantics | 2023 | Weakly Supervised | Captioning | Detection, Explanation | — |
| Ex-VADEx-VAD — LLM Knowledge Enhancement | 2025 | Weakly Supervised | VLMLLM | Detection, Explanation, Classification | — |
| DAKDDAKD — Knowledge Distillation | 2025 | Weakly Supervised | VLM | Detection | — |
| LEC-VADLEC-VAD — Caption-Guided Semantics | 2025 | Weakly Supervised | VLM | Detection | — |
| LOP-VADLOP-VAD — Prompt-Enhanced MIL | 2025 | Weakly Supervised | VLM | Detection | — |
| TCRFLTCRFL — Caption-Guided Semantics | 2025 | Weakly Supervised | VLM | Detection | — |
| FSA-VADFSA-VAD — Caption-Guided Semantics | 2025 | Weakly Supervised | VLM | Detection | — |
| RelVidRelVid — Relation-Aware | 2025 | Weakly Supervised | CLIP | Detection | — |
| AVadCLIPAVadCLIP — Multimodal Fusion | 2025 | Weakly Supervised | CLIPAudio | Detection | — |
| Multimodal-VADMultimodal-VAD — Multimodal Fusion | 2025 | Weakly Supervised | VLMAudio | Detection | — |
| Federated-WVADFederated-WVAD — Prompt-Enhanced MIL | 2025 | Weakly Supervised | VLM | Detection | — |
| IELD-WVADIELD-WVAD — LLM Knowledge Enhancement | 2024 | Weakly Supervised | VLM | Detection, Explanation | — |
| DWFF-VADDWFF-VAD — Multimodal Fusion | 2024 | Weakly Supervised | VLM | Detection | — |
| CMSILCMSIL — Vision-Language Adaptation | 2024 | Weakly Supervised | CLIP | Detection | — |
| WSVAD-LLMKEWSVAD-LLMKE — LLM Knowledge Enhancement | 2024 | Weakly Supervised | LLM | Detection | — |
| CALLMCALLM — LLM Knowledge Enhancement | 2024 | Weakly Supervised | LLM | Detection, Explanation | — |
| MissionGNNMissionGNN — Relation-Aware | 2025 | Weakly Supervised | LLMGNN | Detection, Classification | — |
| ALANALAN — Anomaly Retrieval | 2024 | Weakly Supervised | VLM | Retrieval, Grounding | — |
| LAVADLAVAD — Caption-Based Temporal Reasoning | 2024 | Training-Free | VLMLLM | Detection, Temporal localization, Explanation | — |
| VERAVERA — Verbalized Reasoning | 2025 | Training-Free | VLM | Detection, Explanation | — |
| SUVADSUVAD — Semantic Understanding | 2025 | Training-Free | MLLM | Detection, Explanation | — |
| EventVADEventVAD — Event-Aware Reasoning | 2025 | Training-Free | MLLM | Detection, Temporal localization | — |
| MCANetMCANet — Multimodal Caption-Aware | 2024 | Training-Free | LLMAudio | Detection | — |
| FlashbackFlashback — Memory-Driven | 2025 | Training-Free | LLMCross-modal | Detection, Explanation | — |
| PANDAPANDA — Agentic Pipelines | 2025 | Training-Free | MLLMAgent | Detection, Explanation | — |
| QVADQVAD — Agentic Pipelines | 2026 | Training-Free | VLMAgent | Detection, Retrieval, Grounding | — |
| VADTreeVADTree — Hierarchical Reasoning | 2025 | Training-Free | MLLM | Detection, Explanation | — |
| MoniTorMoniTor — Online Reasoning | 2025 | Training-Free | LLM | Detection | — |
| TFPLGTFPLG — Pseudo-Label Generation | 2025 | Training-Free | VLM | Detection | — |
| AnyAnomalyAnyAnomaly — Customizable Reasoning | 2026 | Training-Free | LVLM | Detection, Question answering | — |
| Anomaly-OVAnomaly-OV — Open-Ended Reasoning | 2025 | Training-Free | MLLM | Detection, Explanation | — |
| VisionGPTVisionGPT — Application-Specific | 2024 | Training-Free | LLM | Detection | — |
| Holmes-VADHolmes-VAD — Anomaly Understanding | 2024 | Instruction-Tuned | MLLM | Detection, Explanation, Question answering | — |
| Holmes-VAUHolmes-VAU — Any-Granularity Understanding | 2025 | Instruction-Tuned | MLLM | Detection, Explanation | — |
| HAWKHAWK — Open-World Understanding | 2024 | Instruction-Tuned | VLM | Explanation, Question answering | — |
| AssistPDAAssistPDA — Online Assistance | 2025 | Instruction-Tuned | MLLM | Detection, Explanation, Question answering | — |
| Vad-R1Vad-R1 — Anomaly Reasoning | 2026 | Instruction-Tuned | MLLM | Detection, Reasoning, Explanation | — |
| VAU-R1VAU-R1 — Anomaly Reasoning | 2025 | Instruction-Tuned | MLLM | Question answering, Grounding, Reasoning | — |
| SherlockSherlock — Structured Extraction | 2025 | Instruction-Tuned | LLM | Detection, Grounding, Classification | — |
| OVVADOVVAD — Detect + Categorize | 2024 | Open-World / Open-Vocabulary | CLIPLLM | Detection, Classification | — |
| AnomizeAnomize — Novel-Category Alignment | 2025 | Open-World / Open-Vocabulary | VLM | Detection, Classification | — |
| PLOVADPLOVAD — Prompt Adaptation | 2025 | Open-World / Open-Vocabulary | VLM | Detection, Classification | — |
| MEL-VLPMEL-VLP — Uncertainty Calibration | 2025 | Open-World / Open-Vocabulary | CLIP | Detection | — |
| LaGoVADLaGoVAD — Language-Guided Definitions | 2025 | Open-World / Open-Vocabulary | VLM | Detection, Classification, Grounding | — |
Classical non-LM baseline. Detects anomalies via large future-frame prediction errors on events that cannot be forecast from normal training video.
Classical non-LM baseline. Constrains reconstruction with a memory module of prototypical normal patterns to avoid reconstructing anomalies too well.
Hybrid SSL-LLM model for complex interaction anomalies: a self-supervised detector selects suspicious frames, then a VLM/LLM validation stage uses refined captions and environment rules to decide whether the event is anomalous.
Uses MLLM-generated descriptions of single-object activities and object-pair interactions as scene-specific normality exemplars; at inference, new descriptions are compared to compact nominal exemplar sets, yielding anomaly scores plus textual evidence.
SlowFastVAD integrates a fast autoencoder-based anomaly detector with a slower RAG-enhanced vision-language model: entropy-based uncertainty selects ambiguous segments for VLM re-analysis using a knowledge base of normal and inferred abnormal patterns, then fuses both detectors' confidence scores.
Fuses appearance, motion, and language-derived semantic cues in a tri-modal encoder; a multimodal decoder predicts next frame and semantic representation, and a Sparse Feature Filtering Module limits over-generalization to abnormal events.
VLAVAD pairs a cross-modal pretrained model with a Selective-Prompt Adapter that selects relevant semantic space, and a Sequence State Space Module that detects temporal inconsistencies in mapped low-dimensional semantic features, enhancing interpretability of unsupervised anomaly detection.
AnomalyRuler is a rule-based LLM framework for one-class VAD: an induction stage derives normality and anomaly rules from few-shot normal reference frames, and a deduction stage applies them to test frames, aided by rule aggregation, perception smoothing, and robust reasoning.
Compares frames or object crops against predefined natural-language descriptions of normal and abnormal situations generated with ChatGPT, learning a text-conditional similarity function without frame- or video-level labels.
Seminal non-LM baseline formulating weakly supervised VAD as multiple instance learning over bags of temporal instances.
Non-LM baseline: transformer-based network using multi-sequence learning, where sequences of multiple snippets serve as the optimization unit instead of single snippets, with a self-training strategy that progressively shortens sequences to refine snippet-level anomaly scores.
Non-LM baseline: multiple instance self-training framework combining a pseudo-label generator with sparse continuous sampling for reliable clip-level labels and a self-guided attention encoder that focuses on anomalous regions to extract task-specific representations.
Non-LM baseline: graph convolutional network that cleans noisy video-level labels by propagating supervision across snippets via a feature-similarity graph and a temporal-consistency graph, turning weakly supervised anomaly detection into a trainable action-classification task.
VadCLIP adapts frozen CLIP to weakly supervised VAD via a dual-branch design: one branch performs coarse-grained binary classification from visual features, the other exploits fine-grained vision-language alignment, achieving state-of-the-art results on UCF-Crime and XD-Violence.
Extends VadCLIP with dual learnable text-prompt branches: a dynamic branch supervises inter-frame difference features to capture temporal changes in anomalous behavior, while a static branch optimizes in-frame localization, jointly improving weakly supervised detection.
Uses ViT-encoded CLIP visual features in place of C3D or I3D, then models temporal dependencies with a Temporal Self-Attention module to nominate snippets of interest for weakly supervised anomaly localization.
Generates reparameterized learnable prompts instead of hand-crafted templates for richer anomaly semantics, then combines a classification block with a video-text alignment block to detect anomalies at both coarse video-level and fine frame-level granularity.
Combines a Temporal Context Aggregation module with a Prompt-Enhanced Learning module that builds knowledge-based prompts from external commonsense concepts to improve fine-grained discriminability among anomaly categories.
TEVAD (Text-Empowered VAD) generates dense captions for video snippets and fuses text and visual features for anomaly detection, improving accuracy and robustness on ShanghaiTech, UCF-Crime, XD-Violence, and UCSD-Pedestrians while using word-level caption contributions to explain predictions.
Ex-VAD combines fine-grained anomaly classification with explanations: a VLM extracts frame-level captions that an LLM converts into video-level explanations, and a Label Augment and Alignment Module expands anomaly category labels into descriptive phrases for more precise detection.
Distills knowledge from aggregated representations of multiple backbones into a single-backbone student model via bi-level distillation and a disentangled cross-attention feature-aggregation network, improving over prior methods on UCF-Crime, ShanghaiTech, and XD-Violence.
Addresses incomplete event localization with a dual vision-language structure encoding category-aware and category-agnostic semantics; an anomaly-aware Gaussian mixture learns precise event boundaries while a memory-bank prototype mechanism enriches anomaly-category text descriptions.
Introduces dual affirmative/negative ("yes/no") text prompts for bidirectional vision-language alignment, addressing overlap between normal and abnormal visual patterns, paired with a temporal context attention module that captures dependencies CLIP's static representations miss.
Combines Mamba and Transformer layers in a temporal context learning module to capture short- and long-range event dependencies with linear complexity, alongside representative feature learning that reduces reliance on noisy top-scoring snippet selection.
Filter, summarize, align: semantic-guided weakly supervised VAD.
Relational learning with vision-language models plus auxiliary anomaly-recognition and reconstruction objectives to improve class separation and temporal feature learning.
AVadCLIP is a weakly supervised audio-visual VAD framework built on frozen CLIP: an efficient audio-visual fusion module adapts cross-modal features, an audio-visual prompt enhances text embeddings, and uncertainty-driven distillation synthesizes audio-visual representations from visual-only input at inference.
Dual-stream multimodal VAD network combining video, audio, and text: a coarse-grained stream cross-modally fuses audio with temporally modeled visual features via contrastive optimization, while a fine-grained stream builds abnormal-aware context prompts to discriminate anomalies via a 'coarse-support-fine' strategy.
Federated weakly supervised VAD that replaces manual text prompts with a text-prompt generator conditioned on global anomaly-related context and personalized local visual context, letting clients train on private surveillance data without sharing raw video.
Proposes TCVADS, a two-stage cross-modal system for edge devices: a distilled convolutional network performs fast coarse-grained classification, then a CLIP-based fine-grained classifier with cross-modal contrastive learning is triggered only when an anomaly is flagged.
Weakly supervised VAD using CLIP to align video and text representations, with dynamic weighting that emphasizes abnormal frames during training and a fusion of video-level and frame-level features to address the mix of normal and abnormal content within anomalous videos.
Two-branch framework pairing a CLIP-driven vision-language branch, which generates pseudo anomaly cues from visual-concept priors, with multi-scale instance learning that models temporal correlation between adjacent snippets to reduce the false alarms typical of standard MIL.
MIL detection framework that queries large language models for anomaly context-boundary prompts to fix the fuzzy temporal boundaries of binary-labeled anomalies, using Anomaly Temporal Feature Amplification and Anomaly Knowledge Enhancement modules that mutually reinforce each other.
CALLM cascades a 3D deep autoencoder with a large video-language model: the autoencoder's binary classification flags abnormality and triggers a follow-up query to the LVLM for detailed explanation, using weakly supervised caption extraction and pseudo-instruction injection to improve performance.
Generates mission-specific knowledge graphs with an LLM and ConceptNet, then applies hierarchical GNNs for weakly supervised anomaly recognition.
ALAN (Anomaly-Led Alignment Network) tackles Video Anomaly Retrieval, retrieving long untrimmed anomalous videos via text or audio queries; it uses anomaly-led sampling on key segments, a pretext task for fine-grained video-text association, and complementary cross-modal alignments, on new UCFCrime-AR and XDViolence-AR benchmarks.
LAVAD is a training-free method that turns VAD into language-based reasoning: a VLM captioning model describes each frame, an LLM aggregates captions over temporal windows to estimate anomaly scores, and cross-modal similarity cleans noisy captions and refines the scores.
VERA is a verbalized learning framework letting frozen VLMs perform explainable VAD without parameter updates: it decomposes VAD reasoning into learnable guiding questions optimized via data-driven verbal interactions on coarsely labeled data, then fuses scene and temporal context into frame-level anomaly scores.
SUVAD is a training-free MLLM-based method that generates detailed textual descriptions of video content for semantic understanding, then detects anomalies directly via an LLM from those descriptions, with techniques to mitigate MLLM hallucination and support flexible, scene-adaptable anomaly definitions.
Segments long videos into event-consistent units via dynamic spatiotemporal graph modeling, statistical boundary detection, and hierarchical prompting before anomaly scoring.
Multimodal caption-aware training-free VAD combining visual, audio, and language cues via a large language model.
Builds an offline pseudo-scene memory of LLM-generated normal and anomalous captions embedded with a frozen encoder; online segments are matched by similarity search, removing expensive online LLM calls for real-time zero-shot VAD.
MLLM-based agentic AI engineer for generalist VAD without training data or human tuning, combining scene-aware RAG strategy planning, heuristic-guided reasoning, tool-augmented self-reflection, and a chain-of-memory mechanism that lets it improve from past cases.
Treats VLM-LLM interaction as a dynamic dialogue instead of static prompting: an LLM agent iteratively refines queries from visual context so lightweight VLMs produce high-fidelity captions, reaching state-of-the-art results with a fraction of the parameters of larger training-free methods.
Builds a Hierarchical Granularity-aware Tree from a pretrained event-boundary detector to replace fixed-length window sampling: generic event nodes are structured coarse-to-fine, scored by VLMs with injected priors and reasoned over by LLMs, then merged via inter-cluster correlation.
First training-free approach to online VAD: streams frames to a VLM while an LSTM-inspired mechanism models past states, and a dynamically updated scoring queue with an anomaly prior guides the LLM to distinguish normal from abnormal behavior over time.
Triple-branch training-free module that uses CLIP's vision-language alignment and a threshold-guided similarity match, rather than a learned classifier, to generate fine- and coarse-grained pseudo-labels, with transformer and GCN branches capturing short- and long-range temporal dependencies.
AnyAnomaly introduces customizable video anomaly detection (C-VAD), where user-defined text specifies the abnormal event; it performs segment-level, context-aware visual question answering with a frozen large vision-language model using key-frame selection and grid-format temporal context, without fine-tuning.
Anomaly-OV (Anomaly-OneVision) is a specialist visual assistant for zero-shot anomaly detection and reasoning with MLLMs, using a Look-Twice Feature Matching mechanism to adaptively select and emphasize abnormal visual tokens; trained on the new Anomaly-Instruct-125k dataset and evaluated on VisA-D&R.
VisionGPT pairs the open-world detector YOLO-World with specialized LLM prompts for zero-shot anomaly detection in navigation frames, generating concise audio-delivered descriptions of obstacles and enabling dynamic scene switching to assist safe visual navigation in complex environments.
Holmes-VAD builds VAD-Instruct50k, a large-scale multimodal instruction-tuning benchmark created via semi-automatic labeling with a video captioner and an LLM, then trains a lightweight temporal sampler to select high-anomaly-response frames and fine-tunes an MLLM for anomaly localization and explanation.
Holmes-VAU introduces HIVAU-70k, a hierarchical clip/event/video-level instruction benchmark built via a semi-automated annotation engine, and an Anomaly-focused Temporal Sampler that combines an anomaly scorer with density-aware sampling to direct a multimodal LLM toward anomaly-rich segments in long videos.
HAWK integrates motion modality into a VLM via an auxiliary consistency loss and explicit motion-to-language supervision, trained on over 8,000 anomaly videos with language descriptions plus 8,000 QA pairs, achieving state-of-the-art video description generation and question-answering performance.
AssistPDA is the first online video surveillance assistant unifying anomaly prediction, detection, and analysis (VAPDA) for streaming video; a Spatio-Temporal Relation Distillation module transfers offline VLM spatiotemporal modeling to real-time inference, trained on the new VAPDA-127K benchmark.
Vad-R1 introduces Video Anomaly Reasoning (VAR), requiring MLLMs to reason step-by-step via a Perception-to-Cognition Chain-of-Thought; trained on the new Vad-Reasoning dataset with supervised fine-tuning followed by AVA-GRPO reinforcement learning using a self-verification reward under limited annotations.
VAU-R1 uses reinforcement fine-tuning on MLLMs to improve anomaly reasoning, decomposing understanding into multiple-choice QA, temporal grounding, and reasoning tasks; paired with VAU-Bench, the first chain-of-thought benchmark for video anomaly reasoning with rationales and temporal annotations.
Sherlock tackles Multi-scene Video Abnormal Event Extraction and Localization (M-VAE): a Global-local Spatial-enhanced MoE module extracts subject/event-type/object/scene quadruples using global-local spatial modeling, while a Spatial Imbalance Regulator balances the mixture-of-experts against skewed spatial-information frequencies.
OVVAD decouples open-vocabulary VAD into class-agnostic detection and class-specific classification, jointly optimized: a semantic knowledge injection module brings LLM knowledge to detection, while an anomaly synthesis module generates pseudo unseen-anomaly videos via large vision generation models for classification.
Anomize targets two challenges in open-vocabulary VAD -- detection ambiguity and categorization confusion -- using multi-level visual data with matching textual information to score novel anomalies, plus label-relation-guided text encoding that better aligns novel videos with their labels.
PLOVAD prompt-tunes frozen image-based vision-language models for open-vocabulary VAD via a Prompting Module -- a learnable domain-specific prompt plus an LLM-crafted anomaly-specific prompt -- and a Temporal Module that stacks a graph attention network atop frame-wise visual features for video-level localization.
Open-world weakly supervised setting using multi-scale temporal modeling and multimodal evidential collaborative learning to estimate uncertainty and dynamically calibrate anomaly boundaries.
LaGoVAD supports open-world VAD with variable anomaly definitions guided by user-provided natural language at inference; it adapts via dynamic video synthesis for duration diversity and contrastive learning with negative mining, trained on the new 35,279-video PreVAD dataset with anomaly descriptions.
Try clearing one or more filters.