Weakly Supervised Video Anomaly Detection with Anomaly-Connected Components and Intention Reasoning
Yu Wang, Shengjie Zhao
Abstract
Weakly supervised video anomaly detection (WS-VAD) involves identifying the temporal intervals that contain anomalous events in untrimmed videos, where only video-level annotations are provided as supervisory signals. However, a key limitation persists in WS-VAD, as dense frame-level annotations are absent, which often leaves existing methods struggling to learn anomaly semantics effectively. To address this issue, we propose a novel framework named LAS-VAD, short for Learning Anomaly Semantics for WS-VAD, which integrates anomaly-connected component mechanism and intention awareness mechanism. The former is designed to assign video frames into distinct semantic groups within a video, and frame segments within the same group are deemed to share identical semantic information. The latter leverages an intention-aware strategy to distinguish between similar normal and abnormal behaviors (e.g., taking items and stealing). To further model the semantic information of anomalies, as anomaly occurrence is accompanied by distinct characteristic attributes (i.e., explosions are characterized by flames and thick smoke), we additionally incorporate anomaly attribute information to guide accurate detection. Extensive experiments on two benchmark datasets, XD-Violence and UCF-Crime, demonstrate that our LAS-VAD outperforms current state-of-the-art methods with remarkable gains.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d782c824-795d-451c-86a2-eaf05f9217f7Builds on29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
- Weakly-supervised Video Anomaly Detection with Robust Temporal Feature Magnitude LearningYu Tian, Guansong Pang, Yuanhong Chen, Rajvinder Singh et al.ICCV 2021 · 495 citations
- Self-Training Multi-Sequence Learning with Transformer for Weakly Supervised Video Anomaly DetectionShuo Li, Fang Liu, Licheng JiaoAAAI 2022 · 282 citations
Related papers
- Learning Event Completeness for Weakly Supervised Video Anomaly DetectionYu Wang, Shiwei ChenICML 2025
- Generalizing Single-Frame Supervision to Event-Level Understanding for Video Anomaly DetectionJunxi Chen, Liang Li, Yunbin Tu, Li Su et al.NeurIPS 2025 · 4 citations
- Mixture of Experts Guided by Gaussian Splatters Matters: A New Approach to Weakly-Supervised Video Anomaly DetectionGiacomo D'Amicantonio, Snehashis Majhi, Quan Kong, Lorenzo Garattoni et al.ICCV 2025 · 5 citations
- TLMA: Mitigating the Impact of Weakly Labeled Information for Video Anomaly DetectionRong Xu, Runqi Wang, Yingjun Zhang, Tao Tao et al.CVPR 2026
- Prompt-Enhanced Multiple Instance Learning for Weakly Supervised Video Anomaly DetectionJunxi Chen, Liang Li, Li Su, Zheng-Jun Zha et al.CVPR 2024
