CMHKF: Cross-Modality Heterogeneous Knowledge Fusion for Weakly Supervised Video Anomaly Detection
Guohua Wang, Shengping Song, Wuchun He, Yongsen Zheng
Abstract
Weakly supervised video anomaly detection (WSVAD) presents a challenging task focused on detecting frame-level anomalies using only video-level labels. However, existing methods focus mainly on visual modalities, neglecting rich multi-modality information. This paper proposes a novel framework, Cross-Modality Heterogeneous Knowledge Fusion (CMHKF), that integrates cross-modality knowledge from video, audio, and text to improve anomaly detection and localization. To achieve adaptive cross-modality heterogeneous knowledge learning, we designed two components: Cross-Modality Video-Text Knowledge Alignment (CVKA) and Audio Modality Feature Adaptive Extraction (AFAE). They extract and aggregate features by exploring inter-modality correlations. By leveraging abundant cross-modality knowledge, our approach improves the discrimination between normal and anomalous segments. Extensive experiments on XD-Violence show our method significantly enhances accuracy and robustness in both coarse-grained and fine-grained anomaly detection.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
- Weakly-supervised Video Anomaly Detection with Robust Temporal Feature Magnitude LearningYu Tian, Guansong Pang, Yuanhong Chen, Rajvinder Singh et al.ICCV 2021 · 495 citations
- Self-Training Multi-Sequence Learning with Transformer for Weakly Supervised Video Anomaly DetectionShuo Li, Fang Liu, Licheng JiaoAAAI 2022 · 282 citations
- MGFN: Magnitude-Contrastive Glance-and-Focus Network for Weakly-Supervised Video Anomaly DetectionYingxian Chen, Zhengzhe Liu, Baoheng Zhang, Wilton W. T. Fok et al.AAAI 2023 · 221 citations
Related papers
- Learning Event Completeness for Weakly Supervised Video Anomaly DetectionYu Wang, Shiwei ChenICML 2025
- TLMA: Mitigating the Impact of Weakly Labeled Information for Video Anomaly DetectionRong Xu, Runqi Wang, Yingjun Zhang, Tao Tao et al.CVPR 2026
- Weakly Supervised Video Anomaly Detection with Anomaly-Connected Components and Intention ReasoningYu Wang, Shengjie ZhaoCVPR 2026 · 6 citations
- Text Prompt with Normality Guidance for Weakly Supervised Video Anomaly DetectionZhiwei Yang, Jing Liu, Peng WuCVPR 2024 · 55 citations
- Prompt-Enhanced Multiple Instance Learning for Weakly Supervised Video Anomaly DetectionJunxi Chen, Liang Li, Li Su, Zheng-Jun Zha et al.CVPR 2024
