Skynet-V1: Towards Early Warning of Video Abnormal Events via A Spatial-temporal Causal-enhanced MoE Framework
Junxiao Ma, Jingjing Wang, Min Zhang, Guodong Zhou
Abstract
In the literature, prior studies on Video Anomaly Detection (VAD) primarily focus on anomalies that have already occurred (i.e., consequential anomaly), but cannot identify the causative anomalies (i.e., the cause of final anomaly), while this type of causative anomaly could be powerfully beneficial to early warning against the anomalies. Meanwhile existing work mainly focuses on classifying whether each video clip is abnormal, and couldn't extract structured video information, such as what is the abnormal type, which people or things are involved, whereas such structured information can potentially contribute to building an efficient system to monitor the above causative and consequential anomalies. To this end, this paper proposes a new chat-paradigm Video Abnormal Events' Early Warning (VAE-EW) task, aiming to localize and extract not only the consequential abnormal event quadruples but also the causative abnormal event quadruples (i.e., subject, predicate, object, and event type). Further, this paper believes that this new task faces two key challenges, i.e., Spatial-temporal modeling challenge and temporal highlighting challenge. On this basis, this paper proposes a new Skynet-V1 Model with a spatial-temporal causal-enhanced Mixture-of-Expert (MoE) Framework, i.e., acting like Skynet in movie 'The Terminator' to track and early warn against abnormal events, for VAE-EW task. Specifically, this model designs a Spatial-temporal Aware MoE Block (SAMB) and a Causal-guided Temporal Enhancing Block (CTEB) to address the two challenges respectively. Extensive experiments on our VAE-EW dataset show the superiority of our model in localizing and extracting abnormal events, especially the causative events, compared to other advanced baseline models, highlighting the importance of the new VAE-EW task and the effectiveness of Skynet-V1 in addressing such task.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get fcd91285-5bd6-466c-9bf3-e2b76f393ff2Related papers
- Sherlock: Towards Multi-scene Video Abnormal Event Extraction and Localization via a Global-local Spatial-sensitive LLMJunxiao Ma, Jingjing Wang, Jiamin Luo, Peiying Yu et al.WWW 2025 · 11 citations
- Weakly Supervised Video Anomaly Detection and Localization with Spatio-Temporal PromptsPeng Wu, Xuerong Zhou, Guansong Pang, Zhiwei Yang et al.ACM MM 2024 · 50 citations
- Mixture of Experts Guided by Gaussian Splatters Matters: A New Approach to Weakly-Supervised Video Anomaly DetectionGiacomo D'Amicantonio, Snehashis Majhi, Quan Kong, Lorenzo Garattoni et al.ICCV 2025 · 5 citations
- Alert-CLIP: Abnormality-aware Latent-Enhanced Representation Tuning of CLIP for Video Anomaly DetectionYiyan Zhu, Menghao Zhang, Haifeng Sun, Pengfei Ren et al.CVPR 2026
- Cloze Test Helps: Effective Video Anomaly Detection via Learning to Complete Video EventsGuang Yu, Siqi Wang, Zhiping Cai, En Zhu et al.ACM MM 2020 · 193 citations
