HiProbe-VAD: Video Anomaly Detection via Hidden States Probing in Tuning-Free Multimodal LLMs
Zhaolin Cai, Fan Li, Ziwei Zheng, Yanjun Qin
Abstract
Video Anomaly Detection (VAD) aims to identify and locate deviations from normal patterns in video sequences. Traditional methods often struggle with substantial computational demands and a reliance on extensive labeled datasets, thereby restricting their practical applicability. To address these constraints, we propose HiProbe-VAD, a novel framework that leverages pre-trained Multimodal Large Language Models (MLLMs) for VAD without requiring fine-tuning. In this paper, we discover that the intermediate hidden states of MLLMs contain information-rich representations, exhibiting higher sensitivity and linear separability for anomalies compared to the output layer. To capitalize on this, we propose a Dynamic Layer Saliency Probing (DLSP) mechanism that intelligently identifies and extracts the most informative hidden states from the optimal intermediate layer during the MLLMs reasoning. Then a lightweight anomaly scorer and temporal localization module efficiently detects anomalies using these extracted hidden states and finally generate explanations. Experiments on the UCF-Crime and XD-Violence datasets demonstrate that HiProbe-VAD outperforms existing training-free and most traditional approaches. Furthermore, our framework exhibits remarkable cross-model generalization capabilities in different MLLMs without any tuning, unlocking the potential of pre-trained MLLMs for video anomaly detection and paving the way for more practical and scalable solutions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f3db28f3-0319-4b35-9c74-114d70d0026cCited by top-tier papers3
- Steering and Rectifying Latent Representation Manifolds in Frozen Multi-modal LLMs for Video Anomaly DetectionZhaolin Cai, Fan Li, Huiyu Duan, Lijun He et al.ICLR 2026 · 2 citations
- HeadHunt-VAD: Hunting Robust Anomaly-Sensitive Heads in MLLM for Tuning-Free Video Anomaly DetectionZhaolin Cai, Fan Li, Ziwei Zheng, Haixia Bi et al.AAAI 2026 · 1 citation
- Linguistic Relative Policy Optimization for Video Anomaly ReasoningJiaxu Leng, Jiankang Zheng, Mengjingcheng Mo, Zhanjie Wu et al.ICML 2026
Builds on29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- Towards Total Recall in Industrial Anomaly DetectionKarsten Roth, Latha Pemula, Joaquin Zepeda, Bernhard Schölkopf et al.CVPR 2022 · 1,301 citations
Related papers
- Harnessing Large Language Models for Training-Free Video Anomaly DetectionLuca Zanella, Willi Menapace, Massimiliano Mancini, Yiming Wang et al.CVPR 2024 · 57 citations
- EventVAD: Training-Free Event-Aware Video Anomaly DetectionYihua Shao, Haojin He, Sijie Li, Siyu Chen et al.ACM MM 2025 · 19 citations
- Ex-VAD: Explainable Fine-grained Video Anomaly Detection Based on Visual-Language ModelsChao Huang, Yushu Shi, Jie Wen, Wei Wang et al.ICML 2025
- Local Patterns Generalize Better for Novel AnomaliesYalong JiangICLR 2025
- MoniTor: Exploiting Large Language Models with Instruction for Online Video Anomaly DetectionShengtian Yang, Yue Feng, Yingshi Liu, Jingrou Zhang et al.NeurIPS 2025 · 16 citations
