Survey Versionv1.0Last UpdatedSeptember 21, 2026Last Changed Inv1.0

Dynamic Survey Maintenance#

Language models have played a dual role in the development of this survey. Their rapid adoption in video anomaly detection has created the need for a focused review of the field, as new LM-based methods, datasets, benchmarks, and evaluation protocols continue to appear at a pace that is difficult to capture through a conventional static survey. At the same time, agentic AI systems have increasingly emerged as a practical paradigm for decomposing complex workflows into specialized machine-learning and maintenance tasks [155]-[159]. The reasoning, retrieval, summarization, and document-editing capabilities of language models therefore provide practical tools for maintaining such a survey over time. We use language models not only as the central subject of this review, but also as components of an agent-assisted maintenance process designed to help the survey keep pace with the literature it studies. In this sense, the same technological developments that motivate the survey also enable its dynamic and versioned form.

To support this goal, we adopt the agentic Dynamic Survey Framework proposed in [36]. The framework treats survey writing as a long-term maintenance problem rather than a one-time document-generation task. Instead of repeatedly producing new surveys with overlapping scope, an existing survey is maintained as a persistent scholarly resource whose content evolves through controlled and versioned revisions. The survey retains an author-defined structure, including its section hierarchy, topical scope, taxonomy, terminology, and table schemas, while newly published work is incorporated incrementally through localized updates rather than global rewriting.

For this survey, the stable structure corresponds to the organization introduced in the preceding sections: problem settings and supervision paradigms, datasets, evaluation metrics, and the taxonomy of LM-based VAD methods, including unsupervised and semi-supervised, weakly supervised, training-free, instruction-tuned, and open-world or open-vocabulary approaches. During maintenance, newly identified papers are first assessed for relevance to language-model-based VAD. This screening step is necessary because not every work involving language models, surveillance video, anomaly terminology, or multimodal reasoning falls within the intended scope. Relevant papers are then routed to the most appropriate section or table according to their primary contribution. For example, a prompt-based zero-shot method may be incorporated into the training-free discussion, a new video-question-answering benchmark may update the dataset section, and a new semantic evaluation protocol may extend the evaluation-metrics discussion.

The update process is deliberately conservative. Once a paper is approved for inclusion, only the smallest relevant portion of the survey should be modified. This may involve adding a new method description, revising an existing comparison, inserting a table entry, updating citation metadata, or adjusting a short transition. Avoiding unnecessary global rewriting helps preserve coherence across versions, including consistent terminology, narrative flow, citation style, and taxonomy. It also reduces the risk that repeated automatic editing gradually changes the survey's scope or weakens the distinctions between method papers, dataset papers, benchmark contributions, and evaluation-oriented work.

The maintenance workflow is agent-assisted but remains human-controlled. Language-model agents can support literature monitoring, technical summarization, relevance filtering, abstention when evidence is insufficient, section routing, table completion, citation placement, and localized text synthesis. However, decisions that affect the high-level scope or structure of the survey remain under explicit author control. In particular, routine updates do not automatically introduce new top-level paradigms, redefine existing categories, or reorganize the taxonomy. Such structural revisions require human review because emerging research trends may initially appear significant but later prove too narrow, temporary, or overlapping with existing categories. More generally, agent-generated outputs are treated as candidate revisions rather than autonomous changes to the scholarly record. Final paper inclusion, categorization, substantive textual revisions, and publication of each maintained release require explicit author review and approval.

Each maintained release is accompanied by a transparent version record. These records identify newly added papers, revised sections, updated tables, corrected metadata, and any author-approved structural changes. Versioning allows readers to distinguish the original peer-reviewed article from later maintained editions while still treating the survey as a coherent scholarly resource. It also supports reproducibility by making the evolution of the document inspectable rather than silently replacing earlier content. If an error is identified in previously released metadata, categorization, or technical description, the correction is documented in the subsequent version record rather than silently overwriting the historical record, while the earlier release remains archived and accessible. Figure 6 summarizes this maintenance process, showing how agent-assisted literature monitoring, relevance screening, section routing, and localized updates are integrated with a persistent survey structure, human approval, and transparent versioned publication.

FIGURE 6. Agent-assisted maintenance workflow for the dynamic LM-based VAD survey. Newly published literature is monitored, screened for relevance, routed to the appropriate survey component, and incorporated through localized updates. The persistent survey core preserves the author-defined taxonomy and structure, while human review precedes publication through the versioned website and archived releases.
FIGURE 6. Agent-assisted maintenance workflow for the dynamic LM-based VAD survey. Newly published literature is monitored, screened for relevance, routed to the appropriate survey component, and incorporated through localized updates. The persistent survey core preserves the author-defined taxonomy and structure, while human review precedes publication through the versioned website and archived releases.

The maintained survey is accompanied by a publicly accessible website at dynamicvadsurvey.github.io, which serves as the reader-facing interface to the dynamic resource, as illustrated in Figure 7. The website provides access to the latest maintained release as well as archived versions, allowing readers to inspect how the survey evolves over time. It follows the same high-level organization as the paper, including supervision paradigms, datasets, evaluation metrics, representative methods, and references. It also supports convenient filtering and exploration of methods according to attributes such as supervision paradigm, publication year, model type, dataset, task, and methodological category, enabling readers to identify and compare relevant approaches without searching through the full document. Version records summarize newly added papers, revised sections, updated tables, and other substantive changes, while a comparison view highlights additions and removals between selected releases.

FIGURE 7. Public web interface of the dynamic LM-based VAD survey. The website provides access to the latest maintained release, archived versions, citation information, version records, and changes introduced across updates.
FIGURE 7. Public web interface of the dynamic LM-based VAD survey. The website provides access to the latest maintained release, archived versions, citation information, version records, and changes introduced across updates.

For the initial implementation of the maintenance pipeline, we favor small and mid-sized language models rather than relying exclusively on the largest available systems. Recent work suggests that smaller models can be effective for specialized agentic tasks when the workflow is decomposed into well-defined components with constrained inputs and outputs [160]-[163]. This design is consistent with the dynamic survey framework, in which discovery, screening, routing, abstention, synthesis, and verification are handled as separate operations rather than as a single unconstrained generation task [36]. The initial system therefore uses models from the Gemma [164], [165] and Qwen [166]-[168] families for lightweight maintenance operations. These backbone choices are not fixed: stronger, more efficient, or more reliable models may be substituted in future releases while preserving the same high-level workflow, governance principles, and versioned update protocol.

Beyond keeping the literature coverage current, the maintained records can reveal broader patterns in the evolution of LM-based VAD, including emerging methodological directions, recurring weaknesses, and areas where benchmark design or evaluation remains inadequate. These observations provide a natural connection to the challenges and future research directions discussed in the next section.