CueBench: Advancing Unified Understanding of Context-Aware Video Anomalies in Real-World
Yating Yu, Congqi Cao, Zhaoying Wang, Weihua Meng, Jie Li, Yuxin Li, Zihao Wei, Zhongpei Shen, Jiajun Zhang
Abstract
How far are deep models from real-world video anomaly understanding (VAU)? Current works typically emphasize detecting unexpected occurrences deviating from normal patterns or comprehending anomalous events with interpretable descriptions. However, they exhibit only a superficial comprehension of real-world anomalies, with limited breadth in complex principles and subtle contexts that distinguish the anomalies from normalities, e.g., climbing cliffs with safety gear vs. without it. To this end, we introduce CueBench, the first of its kind Benchmark, devoted to Context-aware video anomalies within a Unified Evaluation framework. We comprehensively establish an event-centric hierarchical taxonomy that anchors two core event types: 14 conditional and 18 absolute anomaly events, defined by their refined semantics from diverse contexts across 174 scenes and 198 attributes. Based on this, we propose to unify and benchmark context-aware VAU with various challenging tasks across recognition, temporal grounding, detection, and anticipation. It also serves as a rigorous and fair probing evaluation suite for generalized and specialized vision-language models (VLMs) across both generative and discriminative paradigms. To address the challenges underlying CueBench, we further develop Cue-R1 based on R1-style reinforcement fine-tuning with verifiable, task-aligned, and hierarchy-refined rewards in a unified generative manner. Extensive results on CueBench reveal that, existing VLMs are still far from satisfactory real-world anomaly understanding, while our Cue-R1 surpasses these state-of-the-art approaches by over 24% on average.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b43c142f-afae-4540-a18f-3993498a96e4Builds on29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Visual-RFT: Visual Reinforcement Fine-TuningZiyu Liu, Zeyi Sun, Yuhang Zang, Xiaoyi Dong et al.ICCV 2025 · 563 citations
- Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language ModelsMuhammad Maaz, Hanoona Abdul Rasheed, Salman Khan, Fahad KhanACL 2024 · 279 citations
Related papers
- VALU: A Benchmark for Video Anomaly Temporal Localization and Understanding at Multiple Semantic LevelsYixiao He, Menghao Zhang, Haifeng Sun, Jing Wang et al.ACL 2026
- FineVAU: A Novel Human-Aligned Benchmark for Fine-Grained Video Anomaly UnderstandingJoão Alexandre Cardeira Pereira, Vasco Lopes, João Neves, David SemedoAAAI 2026
- VAGU & GtS: LLM-Based Benchmark and Framework for Joint Video Anomaly Grounding and UnderstandingShibo Gao, Peipei Yang, Yangyang Liu, Yi Chen et al.AAAI 2026 · 5 citations
- CAVE : Detecting and Explaining Commonsense Anomalies in Visual EnvironmentsRishika Bhagwatkar, Syrielle Montariol, Angelika Romanou, Beatriz Borges et al.EMNLP 2025
- Uncovering what, why and How: A Comprehensive Benchmark for Causation Understanding of Video AnomalyHang Du, Sicheng Zhang, Binzhu Xie, Guoshun Nan et al.CVPR 2024
