Effective Video Abnormal Event Detection by Learning A Consistency-Aware High-Level Feature Extractor
Guang Yu, Siqi Wang, Zhiping Cai, Xinwang Liu, Chengkun Wu
Abstract
With pure normal training videos, video abnormal event detection (VAD) aims to build a normality model, and then detect abnormal events that deviate from this model. Despite of some progress, existing VAD methods typically train the normality model by a low-level learning objective (e.g. pixel-wise reconstruction/prediction), which often overlooks the high-level semantics in videos. To better exploit high-level semantics for VAD, we propose a novel paradigm that performs VAD by learning a Consistency-Aware high-level Feature Extractor (CAFE). Specifically, with a pre-trained deep neural network (DNN) as teacher network, we first feed raw video events into the teacher network and extract the outputs of multiple hidden layers as their high-level features, which contain rich high-level semantics. Guided by high-level features extracted from normal training videos, we train a student network to be the high-level feature extractor of normal events, so as to explicitly consider high-level semantics in training. For inference, a video event can be viewed as normal if the student extractor produces similar high-level features to the teacher network. Second, based on the fact that consecutive video frames usually enjoy minor differences, we propose a consistency-aware scheme that requires high-level features extracted from neighboring frames to be consistent. Our consistency-aware scheme not only encourages the student extractor to ignore low-level differences and capture more high-level semantics, but also enables better anomaly scoring. Last, we also design a generic framework that can bridge high-level and low-level learning in VAD to further ameliorate VAD performance. By flexibly embedding one or more low-level learning objectives into CAFE, the framework makes it possible to combine the strengths of both high-level and low-level learning. The proposed method attains state-of-the-art results on commonly-used benchmark datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on18
- Memorizing Normality to Detect Anomaly: Memory-Augmented Deep Autoencoder for Unsupervised Anomaly DetectionDong Gong, Lingqiao Liu, Vuong Le, Budhaditya Saha et al.ICCV 2019 · 1,646 citations
- Anomaly Detection in Video Sequence With Appearance-Motion CorrespondenceTrong-Nguyen Nguyen, Jean MeunierICCV 2019 · 414 citations
- A Hybrid Video Anomaly Detection Framework via Memory-Augmented Flow Reconstruction and Flow-Guided Frame PredictionZhian Liu, Yongwei Nie, Chengjiang Long, Qing Zhang et al.ICCV 2021 · 341 citations
- Appearance-Motion Memory Consistency Network for Video Anomaly DetectionRuichu Cai, Hao Zhang, Wen Liu, Shenghua Gao et al.AAAI 2021 · 223 citations
- Cloze Test Helps: Effective Video Anomaly Detection via Learning to Complete Video EventsGuang Yu, Siqi Wang, Zhiping Cai, En Zhu et al.ACM MM 2020 · 193 citations
Related papers
- Learning Causality-inspired Representation Consistency for Video Anomaly DetectionYang Liu, Zhaoyang Xia, Mengyang Zhao, Donglai Wei et al.ACM MM 2023 · 48 citations
- Video Event Restoration Based on Keyframes for Video Anomaly DetectionZhiwei Yang, Jing Liu, Zhaoyang Wu, Peng Wu et al.CVPR 2023
- Fine-VAD: Towards Fine-Grained Video Anomaly Detection via Progressive Cross-Granularity LearningMenghao Zhang, Yiyan Zhu, Pengfei Ren, Haifeng Sun et al.CVPR 2026
- Text Prompt with Normality Guidance for Weakly Supervised Video Anomaly DetectionZhiwei Yang, Jing Liu, Peng WuCVPR 2024 · 55 citations
- Cluster Attention Contrast for Video Anomaly DetectionZiming Wang, Yuexian Zou, Zeming ZhangACM MM 2020 · 90 citations
