RefineVAD: Semantic-Guided Feature Recalibration for Weakly Supervised Video Anomaly Detection
Junhee Lee, ChaeBeen Bang, MyoungChul Kim, MyeongAh Cho
Abstract
Weakly-Supervised Video Anomaly Detection aims to identify anomalous events using only video-level labels, balancing annotation efficiency with practical applicability. However, existing methods often oversimplify the anomaly space by treating all abnormal events as a single category, overlooking the diverse semantic and temporal characteristics intrinsic to real-world anomalies. Inspired by how humans perceive anomalies, by jointly interpreting temporal motion patterns and semantic structures underlying different anomaly types, we propose RefineVAD, a novel framework that mimics this dual-process reasoning. Our framework integrates two core modules. The first, Motion-aware Temporal Attention and Recalibration (MoTAR), estimates motion salience and dynamically adjusts temporal focus via shift-based attention and global Transformer-based modeling. The second, Category-Oriented Refinement (CORE), injects soft anomaly category priors into the representation space by aligning segment-level features with learnable category prototypes through cross-attention. By jointly leveraging temporal dynamics and semantic structure, explicitly models both how'' motion evolves and what'' semantic category it resembles. Extensive experiments on WVAD benchmark validate the effectiveness of RefineVAD and highlight the importance of integrating semantic context to guide feature refinement toward anomaly-relevant patterns.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e2cc460b-6891-4953-a398-ebb1968aad95Builds on15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- Weakly-supervised Video Anomaly Detection with Robust Temporal Feature Magnitude LearningYu Tian, Guansong Pang, Yuanhong Chen, Rajvinder Singh et al.ICCV 2021 · 495 citations
- Appearance-Motion Memory Consistency Network for Video Anomaly DetectionRuichu Cai, Hao Zhang, Wen Liu, Shenghua Gao et al.AAAI 2021 · 223 citations
- VadCLIP: Adapting Vision-Language Models for Weakly Supervised Video Anomaly DetectionPeng Wu, Xuerong Zhou, Guansong Pang, Lingru Zhou et al.AAAI 2024 · 220 citations
Related papers
- Fine-VAD: Towards Fine-Grained Video Anomaly Detection via Progressive Cross-Granularity LearningMenghao Zhang, Yiyan Zhu, Pengfei Ren, Haifeng Sun et al.CVPR 2026
- Prompt-Enhanced Multiple Instance Learning for Weakly Supervised Video Anomaly DetectionJunxi Chen, Liang Li, Li Su, Zheng-Jun Zha et al.CVPR 2024
- Learning Event Completeness for Weakly Supervised Video Anomaly DetectionYu Wang, Shiwei ChenICML 2025
- Weakly Supervised Video Anomaly Detection and Localization with Spatio-Temporal PromptsPeng Wu, Xuerong Zhou, Guansong Pang, Zhiwei Yang et al.ACM MM 2024 · 50 citations
- Learning to Tell Apart: Weakly Supervised Video Anomaly Detection via Disentangled Semantic AlignmentWenti Yin, Huaxin Zhang, Xiang Wang, Yuqing Lu et al.AAAI 2026
