Harnessing Large Language Models for Training-Free Video Anomaly Detection
Luca Zanella, Willi Menapace, Massimiliano Mancini, Yiming Wang, Elisa Ricci
Abstract
Video anomaly detection (VAD) aims to temporally locate abnormal events in a video. Existing works mostly rely on training deep models to learn the distribution of normality with either video-level supervision, one-class supervision, or in an unsupervised setting. Training-based methods are prone to be domain-specific, thus being costly for practical deployment as any domain change will involve data collection and model training. In this paper, we radically depart from previous efforts and propose LAnguage-based VAD (LAVAD), a method tackling VAD in a novel, training-free paradigm, exploiting the capabilities of pre-trained large language models (LLMs) and existing vision-language models (VLMs). We leverage VLM-based captioning models to generate textual descriptions for each frame of any test video. With the textual scene description, we then devise a prompting mechanism to unlock the capability of LLMs in terms of temporal aggregation and anomaly score estimation, turning LLMs into an effective video anomaly detector. We further leverage modality-aligned VLMs and propose effective techniques based on cross-modal similarity for cleaning noisy captions and refining the LLM-based anomaly scores. We evaluate LAVAD on two large datasets featuring real-world surveillance scenarios (UCF-Crime and XD- Violence), showing that it outperforms both unsupervised and one-class methods without requiring any training or data collection.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 34611020-90b1-47cc-83ce-4fe23095d298Cited by top-tier papers37
- Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-ThoughtChao Huang, Benfeng Wang, Wei Wang, Jie Wen et al.NeurIPS 2025 · 30 citations
- VADTree: Explainable Training-Free Video Anomaly Detection via Hierarchical Granularity-Aware TreeWenlong Li, Yifei Xu, Yuan Rao, Zhenhua Wang et al.NeurIPS 2025 · 26 citations
- PANDA: Towards Generalist Video Anomaly Detection via Agentic AI EngineerZhiwei Yang, Chen Gao, Mike Zheng ShouNeurIPS 2025 · 24 citations
- EventVAD: Training-Free Event-Aware Video Anomaly DetectionYihua Shao, Haojin He, Sijie Li, Siyu Chen et al.ACM MM 2025 · 19 citations
- MoniTor: Exploiting Large Language Models with Instruction for Online Video Anomaly DetectionShengtian Yang, Yue Feng, Yingshi Liu, Jingrou Zhang et al.NeurIPS 2025 · 16 citations
Builds on16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Weakly-supervised Video Anomaly Detection with Robust Temporal Feature Magnitude LearningYu Tian, Guansong Pang, Yuanhong Chen, Rajvinder Singh et al.ICCV 2021 · 495 citations
- A Hybrid Video Anomaly Detection Framework via Memory-Augmented Flow Reconstruction and Flow-Guided Frame PredictionZhian Liu, Yongwei Nie, Chengjiang Long, Qing Zhang et al.ICCV 2021 · 341 citations
- AnomalyGPT: Detecting Industrial Anomalies Using Large Vision-Language ModelsZhaopeng Gu, Bingke Zhu, Guibo Zhu, Yingying Chen et al.AAAI 2024 · 312 citations
Related papers
- Local Patterns Generalize Better for Novel AnomaliesYalong JiangICLR 2025
- TD-VAD: Breaking Visual Dependence in Video Anomaly Detection with Text-Driven LearningShuangqing Zhang, Lei-Lei Ma, Zhao Wang, Wen Dong et al.ICML 2026
- Ex-VAD: Explainable Fine-grained Video Anomaly Detection Based on Visual-Language ModelsChao Huang, Yushu Shi, Jie Wen, Wei Wang et al.ICML 2025
- HiProbe-VAD: Video Anomaly Detection via Hidden States Probing in Tuning-Free Multimodal LLMsZhaolin Cai, Fan Li, Ziwei Zheng, Yanjun QinACM MM 2025 · 4 citations
- Weakly Supervised Video Anomaly Detection and Localization with Spatio-Temporal PromptsPeng Wu, Xuerong Zhou, Guansong Pang, Zhiwei Yang et al.ACM MM 2024 · 50 citations
