HAWK: Learning to Understand Open-World Video Anomalies
Jiaqi Tang, Hao Lu, Ruizheng Wu, Xiaogang Xu, Ke Ma, Cheng Fang, Bin Guo, Jiangbo Lu, Qifeng Chen, Yingcong Chen
Abstract
Video Anomaly Detection (VAD) systems can autonomously monitor and identify disturbances, reducing the need for manual labor and associated costs. However, current VAD systems are often limited by their superficial semantic understanding of scenes and minimal user interaction. Additionally, the prevalent data scarcity in existing datasets restricts their applicability in open-world scenarios. In this paper, we introduce Hawk, a novel framework that leverages interactive large Visual Language Models (VLM) to interpret video anomalies precisely. Recognizing the difference in motion information between abnormal and normal videos, Hawk explicitly integrates motion modality to enhance anomaly identification. To reinforce motion attention, we construct an auxiliary consistency loss within the motion and video space, guiding the video branch to focus on the motion modality. Moreover, to improve the interpretation of motion-to-language, we establish a clear supervisory relationship between motion and its linguistic representation. Furthermore, we have annotated over 8,000 anomaly videos with language descriptions, enabling effective training across diverse open-world scenarios, and also created 8,000 question-answering pairs for users' open-world questions. The final results demonstrate that Hawk achieves SOTA performance, surpassing existing baselines in both video description generation and question-answering. Our codes/dataset/demo will be released at https://github.com/jqtangust/hawk.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- VAGU & GtS: LLM-Based Benchmark and Framework for Joint Video Anomaly Grounding and UnderstandingShibo Gao, Peipei Yang, Yangyang Liu, Yi Chen et al.AAAI 2026 · 5 citations
- Language-guided Open-world Video Anomaly Detection under Weak SupervisionZihao Liu, Xiaoyu Wu, Jianqin Wu, Xuxu Wang et al.ICLR 2026 · 5 citations
- Do LVLMs Truly Understand Video Anomalies? Revealing Hallucination via Co-Occurrence PatternsMenghao Zhang, Huazheng Wang, Pengfei Ren, Kangheng Lin et al.NeurIPS 2025 · 4 citations
- CueBench: Advancing Unified Understanding of Context-Aware Video Anomalies in Real-WorldYating Yu, Congqi Cao, Zhaoying Wang, Weihua Meng et al.AAAI 2026 · 1 citation
- FineVAU: A Novel Human-Aligned Benchmark for Fine-Grained Video Anomaly UnderstandingJoão Alexandre Cardeira Pereira, Vasco Lopes, João Neves, David SemedoAAAI 2026
Builds on14
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- Frozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalMax Bain, Arsha Nagrani, Gül Varol, Andrew ZissermanICCV 2021 · 1,550 citations
- Weakly-supervised Video Anomaly Detection with Robust Temporal Feature Magnitude LearningYu Tian, Guansong Pang, Yuanhong Chen, Rajvinder Singh et al.ICCV 2021 · 495 citations
Related papers
- Ex-VAD: Explainable Fine-grained Video Anomaly Detection Based on Visual-Language ModelsChao Huang, Yushu Shi, Jie Wen, Wei Wang et al.ICML 2025
- Harnessing Large Language Models for Training-Free Video Anomaly DetectionLuca Zanella, Willi Menapace, Massimiliano Mancini, Yiming Wang et al.CVPR 2024 · 57 citations
- No Need For Real Anomaly: MLLM Empowered Zero-Shot Video Anomaly DetectionZunkai Dai, Ke Li, Jiajia Liu, Jie Yang et al.CVPR 2026 · 6 citations
- Local Patterns Generalize Better for Novel AnomaliesYalong JiangICLR 2025
- HeadHunt-VAD: Hunting Robust Anomaly-Sensitive Heads in MLLM for Tuning-Free Video Anomaly DetectionZhaolin Cai, Fan Li, Ziwei Zheng, Haixia Bi et al.AAAI 2026 · 1 citation
