Multimodal Motion Conditioned Diffusion Model for Skeleton-based Video Anomaly Detection
Alessandro Flaborea, Luca Collorone, Guido Maria D'Amely di Melendugno, Stefano D'Arrigo, Bardh Prenkaj, Fabio Galasso
Abstract
Anomalies are rare and anomaly detection is often therefore framed as One-Class Classification (OCC), i.e. trained solely on normalcy. Leading OCC techniques constrain the latent representations of normal 1 motions to limited volumes and detect as abnormal anything outside, which accounts satisfactorily for the openset'ness of anomalies. But normalcy shares the same openset'ness property since humans can perform the same action in several ways, which the leading techniques neglect. We propose a novel generative model for video anomaly detection (VAD), which assumes that both normality and abnormality are multimodal. We consider skeletal representations and leverage state-of-the-art diffusion probabilistic models to generate multimodal future human poses. We contribute a novel conditioning on the past motion of people and exploit the improved mode coverage capabilities of diffusion processes to generate different-but-plausible future motions. Upon the statistical aggregation of future modes, an anomaly is detected when the generated set of motions is not pertinent to the actual future. We validate our model on 4 established benchmarks: UBnormal, HR-UBnormal, HR-STC, and HR-Avenue, with extensive experiments surpassing state-of-the-art results.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 246de447-216e-48e4-8f92-af794fe02a2aCited by top-tier papers13
- On Diffusion Modeling for Anomaly DetectionVictor Livernoche, Vineet Jain, Yashar Hezaveh, Siamak RavanbakhshICLR 2024 · 74 citations
- Multi-Scale Video Anomaly Detection by Multi-Grained Spatio-Temporal Representation LearningMenghao Zhang, Jingyu Wang, Qi Qi, Haifeng Sun et al.CVPR 2024 · 29 citations
- MULDE: Multiscale Log-Density Estimation via Denoising Score Matching for Video Anomaly DetectionJakub Micorek, Horst Possegger, Dominik Narnhofer, Horst Bischof et al.CVPR 2024 · 21 citations
- Dual Conditioned Motion Diffusion for Pose-Based Video Anomaly DetectionHongsong Wang, Andi Xu, Pinle Ding, Jie GuiAAAI 2025 · 8 citations
- EasyTune: Efficient Step-Aware Fine-Tuning for Diffusion-Based Motion GenerationXiaofeng Tan, Wanjiang Weng, Haodong Lei, Hongsong WangICLR 2026 · 6 citations
Builds on15
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 5,234 citations
Related papers
- UBnormal: New Benchmark for Supervised Open-Set Video Anomaly DetectionAndra Acsintoae, Andrei Florescu, Mariana-Iuliana Georgescu, Tudor Mare et al.CVPR 2022 · 153 citations
- Normalizing Flows for Human Pose Anomaly DetectionOr Hirschorn, Shai AvidanICCV 2023 · 97 citations
- Video Anomaly Detection with Motion and Appearance Guided Patch Diffusion ModelHang Zhou, Jiale Cai, Yuteng Ye, Yonghui Feng et al.AAAI 2025 · 23 citations
- Graph Embedded Pose Clustering for Anomaly DetectionAmir Markovitz, Gilad Sharir, Itamar Friedman, Lihi Zelnik-Manor et al.CVPR 2020
- A Multilevel Guidance-Exploration Network and Behavior-Scene Matching Method for Human Behavior Anomaly DetectionGuoqing Yang, Zhiming Luo, Jianzhe Gao, Yingxin Lai et al.ACM MM 2024 · 1 citation
