Convolutional Transformer based Dual Discriminator Generative Adversarial Networks for Video Anomaly Detection
Xinyang Feng, Dongjin Song, Yuncong Chen, Zhengzhang Chen, Jingchao Ni, Haifeng Chen
Abstract
Detecting abnormal activities in real-world surveillance videos is an important yet challenging task as the prior knowledge about video anomalies is usually limited or unavailable. Despite that many approaches have been developed to resolve this problem, few of them can capture the normal spatio-temporal patterns effectively and efficiently. Moreover, existing works seldom explicitly consider the local consistency at frame level and global coherence of temporal dynamics in video sequences. To this end, we propose Convolutional Transformer based Dual Discriminator Generative Adversarial Networks (CT-D2GAN) to perform unsupervised video anomaly detection. Specifically, we first present a convolutional transformer to perform future frame prediction. It contains three key components, i.e., a convolutional encoder to capture the spatial information of the input video clips, a temporal self-attention module to encode the temporal dynamics, and a convolutional decoder to integrate spatio-temporal features and predict the future frame. Next, a dual discriminator based adversarial training procedure, which jointly considers an image discriminator that can maintain the local consistency at frame-level and a video discriminator that can enforce the global coherence of temporal dynamics, is employed to enhance the future frame prediction. Finally, the prediction error is used to identify abnormal video frames. Thoroughly empirical studies on three public video anomaly detection datasets, i.e., UCSD Ped2, CUHK Avenue, and Shanghai Tech Campus, demonstrate the effectiveness of the proposed adversarial spatio-temporal modeling framework.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cb3363eb-df1d-4260-a4b5-d64479a42365Cited by top-tier papers12
- Feature Prediction Diffusion Model for Video Anomaly DetectionCheng Yan, Shiyu Zhang, Yang Liu, Guansong Pang et al.ICCV 2023 · 76 citations
- Weakly Supervised Video Anomaly Detection and Localization with Spatio-Temporal PromptsPeng Wu, Xuerong Zhou, Guansong Pang, Zhiwei Yang et al.ACM MM 2024 · 50 citations
- Deep Anomaly Discovery from Unlabeled Videos via Normality Advantage and Self-Paced RefinementGuang Yu, Siqi Wang, Zhiping Cai, Xinwang Liu et al.CVPR 2022 · 37 citations
- Evidential Reasoning for Video Anomaly DetectionChe Sun, Yunde Jia, Yuwei WuACM MM 2022 · 23 citations
- MULDE: Multiscale Log-Density Estimation via Denoising Score Matching for Video Anomaly DetectionJakub Micorek, Horst Possegger, Dominik Narnhofer, Horst Bischof et al.CVPR 2024 · 21 citations
Builds on2
Related papers
- Dual Conditioned Motion Diffusion for Pose-Based Video Anomaly DetectionHongsong Wang, Andi Xu, Pinle Ding, Jie GuiAAAI 2025 · 8 citations
- Dual Memory Units with Uncertainty Regulation for Weakly Supervised Video Anomaly DetectionHang Zhou, Junqing Yu, Wei YangAAAI 2023 · 180 citations
- Frequency-Aware Spatiotemporal Transformers for Video Inpainting DetectionBingyao Yu, Wanhua Li, Xiu Li, Jiwen Lu et al.ICCV 2021 · 38 citations
- Self-Supervised Video GANs: Learning for Appearance Consistency and Motion CoherencySangeek Hyun, Jihwan Kim, Jae-Pil HeoCVPR 2021
- Video Event Restoration Based on Keyframes for Video Anomaly DetectionZhiwei Yang, Jing Liu, Zhaoyang Wu, Peng Wu et al.CVPR 2023
