Federated Weakly Supervised Video Anomaly Detection with Multimodal Prompt
Benfeng Wang, Chao Huang, Jie Wen, Wei Wang, Yabo Liu, Yong Xu
Abstract
Video anomaly detection (VAD) aims at locating the abnormal events in videos. Recently, the Weakly Supervised VAD has made great progress, which only requires video-level annotations when training. In practical applications, different institutions may have different types of abnormal videos. However, the abnormal videos cannot be circulated on the internet due to privacy protection. To train a more generalized anomaly detector that can identify various anomalies, it is reasonable to introduce federated learning into WSVAD. In this paper, we propose Global and Local Context-driven Federated Learning, a new paradigm for privacy protected weakly supervised video anomaly detection. Specifically, we utilize the vision-language association of CLIP to detect whether the video frame is abnormal. Instead of leveraging handcrafted text prompts for CLIP, we propose a text prompt generator. The generated prompt is simultaneously influenced by text and visual. On the one hand, the text provides global context related to anomaly, which improves the model's ability of generalization. On the other hand, the visual provides personalized local context because different clients may have videos with different types of anomalies or scenes. The generated prompt ensures global generalization while processing personalized data from different clients. Extensive experiments show that the proposed method achieves remarkable performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 52e76147-8b30-4f61-a047-da4d2c0aa99aCited by top-tier papers7
- Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-ThoughtChao Huang, Benfeng Wang, Wei Wang, Jie Wen et al.NeurIPS 2025 · 30 citations
- Weakly Supervised Video Anomaly Detection with Anomaly-Connected Components and Intention ReasoningYu Wang, Shengjie ZhaoCVPR 2026 · 6 citations
- Enhancing Visual Representation with Textual Semantics: Textual Semantics-Powered Prototypes for Heterogeneous Federated LearningXinghao Wu, Jianwei Niu, Xuefeng Liu, Guogang Zhu et al.CVPR 2026 · 4 citations
- TD-VAD: Breaking Visual Dependence in Video Anomaly Detection with Text-Driven LearningShuangqing Zhang, Lei-Lei Ma, Zhao Wang, Wen Dong et al.ICML 2026
- Fine-VAD: Towards Fine-Grained Video Anomaly Detection via Progressive Cross-Granularity LearningMenghao Zhang, Yiyan Zhu, Pengfei Ren, Haifeng Sun et al.CVPR 2026
Builds on19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Learning to Prompt for Open-Vocabulary Object Detection with Vision-Language ModelYu Du, Fangyun Wei, Zihe Zhang, Miaojing Shi et al.CVPR 2022 · 311 citations
- VadCLIP: Adapting Vision-Language Models for Weakly Supervised Video Anomaly DetectionPeng Wu, Xuerong Zhou, Guansong Pang, Lingru Zhou et al.AAAI 2024 · 220 citations
- Dual Memory Units with Uncertainty Regulation for Weakly Supervised Video Anomaly DetectionHang Zhou, Junqing Yu, Wei YangAAAI 2023 · 180 citations
- Revisiting Classifier: Transferring Vision-Language Models for Video RecognitionWenhao Wu, Zhun Sun, Wanli OuyangAAAI 2023 · 141 citations
Related papers
- Weakly Supervised Video Anomaly Detection and Localization with Spatio-Temporal PromptsPeng Wu, Xuerong Zhou, Guansong Pang, Zhiwei Yang et al.ACM MM 2024 · 50 citations
- Prompt-Enhanced Multiple Instance Learning for Weakly Supervised Video Anomaly DetectionJunxi Chen, Liang Li, Li Su, Zheng-Jun Zha et al.CVPR 2024
- Text Prompt with Normality Guidance for Weakly Supervised Video Anomaly DetectionZhiwei Yang, Jing Liu, Peng WuCVPR 2024 · 55 citations
- Learning Event Completeness for Weakly Supervised Video Anomaly DetectionYu Wang, Shiwei ChenICML 2025
- CMHKF: Cross-Modality Heterogeneous Knowledge Fusion for Weakly Supervised Video Anomaly DetectionGuohua Wang, Shengping Song, Wuchun He, Yongsen ZhengACL 2025 · 2 citations
