Modality-aware Contrastive Instance Learning with Self-Distillation for Weakly-Supervised Audio-Visual Violence Detection
Jiashuo Yu, Jinyu Liu, Ying Cheng, Rui Feng, Yuejie Zhang
摘要
Weakly-supervised audio-visual violence detection aims to distinguish snippets containing multimodal violence events with video-level labels. Many prior works perform audio-visual integration and interaction in an early or intermediate manner, yet overlooking the modality heterogeneousness over the weakly-supervised setting. In this paper, we analyze the modality asynchrony and undifferentiated instances phenomena of the multiple instance learning (MIL) procedure, and further investigate its negative impact on weakly-supervised audio-visual learning. To address these issues, we propose a modality-aware contrastive instance learning with self-distillation (MACIL-SD) strategy . Specifically, we leverage a lightweight two-stream network to generate audio and visual bags, in which unimodal background, violent, and normal instances are clustered into semi-bags in an unsupervised way. Then audio and visual violent semi-bag representations are assembled as positive pairs, and violent semi-bags are combined with background and normal instances in the opposite modality as contrastive negative pairs. Furthermore, a self-distillation module is applied to transfer unimodal visual knowledge to the audio-visual model, which alleviates noises and closes the semantic gap between unimodal and multimodal features. Experiments show that our framework outperforms previous methods with lower complexity on the large-scale XD-Violence dataset. Results also demonstrate that our proposed approach can be used as plug-in modules to enhance other networks. Codes are available at https://github.com/JustinYuu/MACIL_SD.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- VADTree: Explainable Training-Free Video Anomaly Detection via Hierarchical Granularity-Aware TreeWenlong Li, Yifei Xu, Yuan Rao, Zhenhua Wang 等NeurIPS 2025 · 被引用 26 次
- Beyond Euclidean: Dual-Space Representation Learning for Weakly Supervised Video Violence DetectionJiaxu Leng, Zhanjie Wu, Mingpi Tan, Yiran Liu 等NeurIPS 2024 · 被引用 22 次
- Weakly Supervised Video Anomaly Detection with Anomaly-Connected Components and Intention ReasoningYu Wang, Shengjie ZhaoCVPR 2026 · 被引用 6 次
- Mixture of Experts Guided by Gaussian Splatters Matters: A New Approach to Weakly-Supervised Video Anomaly DetectionGiacomo D'Amicantonio, Snehashis Majhi, Quan Kong, Lorenzo Garattoni 等ICCV 2025 · 被引用 5 次
- Generalizing Single-Frame Supervision to Event-Level Understanding for Video Anomaly DetectionJunxi Chen, Liang Li, Yunbin Tu, Li Su 等NeurIPS 2025 · 被引用 4 次
它引用的顶会 Paper27
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Big Self-Supervised Models are Strong Semi-Supervised LearnersTing Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi 等NeurIPS 2020 · 被引用 2,611 次
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 被引用 1,305 次
相关 Paper
- SCLAV: Supervised Cross-modal Contrastive Learning for Audio-Visual CodingChao Sun, Min Chen, Jialiang Cheng, Han Liang 等ACM MM 2023 · 被引用 3 次
- TLMA: Mitigating the Impact of Weakly Labeled Information for Video Anomaly DetectionRong Xu, Runqi Wang, Yingjun Zhang, Tao Tao 等CVPR 2026
- XKD: Cross-Modal Knowledge Distillation with Domain Alignment for Video Representation LearningPritam Sarkar, Ali EtemadAAAI 2024 · 被引用 45 次
- Learning to Tell Apart: Weakly Supervised Video Anomaly Detection via Disentangled Semantic AlignmentWenti Yin, Huaxin Zhang, Xiang Wang, Yuqing Lu 等AAAI 2026
- CMHKF: Cross-Modality Heterogeneous Knowledge Fusion for Weakly Supervised Video Anomaly DetectionGuohua Wang, Shengping Song, Wuchun He, Yongsen ZhengACL 2025 · 被引用 2 次
