Learning to Watch: Active Video Anomaly Understanding via Interleaved Policy Optimization
Mengjingcheng Mo, Jiaxu Leng, Xinbo Gao
Abstract
Video anomaly understanding (VAU) relies on sparse, context-dependent cues. However, existing passive paradigms suffer from observational aliasing, where static sampling fails to disambiguate semantically distinct events. To overcome this, we propose , a closed-loop framework that reconceptualizes video understanding as an active sequential decision-making process within a dynamic environment. Inspired by human video-reviewing behavior, this framework unifies internal cognitive reasoning and strategic evidence acquisition into an interleaved policy, utilizing temporal atomic operators such as local backtracking, temporal expansion, and fine-grained sampling to endow the model with perceptual proactivity. To learn such complex interaction strategies under video-level weak supervision, we design Interactive Direct Preference Optimization (iDPO) to achieve trajectory-level policy alignment, guided by an Active Evidence Inquiry (AEI) utility that balances task success, informative evidence acquisition, and interaction cost. This approach enables the agent to learn to actively disambiguate hypotheses while suppressing redundant exploration. Extensive experiments demonstrate that our framework, with only 2B parameters, achieves highly competitive performance, significantly outperforming state-of-the-art large-scale VAU models in complex scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c8a9391c-eb36-482f-b41d-3c847b2c11c9Builds on7
- UBnormal: New Benchmark for Supervised Open-Set Video Anomaly DetectionAndra Acsintoae, Andrei Florescu, Mariana-Iuliana Georgescu, Tudor Mare et al.CVPR 2022 · 153 citations
- Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-ThoughtChao Huang, Benfeng Wang, Wei Wang, Jie Wen et al.NeurIPS 2025 · 30 citations
- JUDO: A Juxtaposed Domain-Oriented Multimodal Reasoner for Industrial Anomaly QAHyunju Kang, Woohyun Lee, Jaewon Kim, Hogun ParkICLR 2026 · 4 citations
- Aligning Effective Tokens with Video Anomaly in Large Language ModelsYingxian Chen, Jiahui Liu, Ruidi Fan, Yanwei Li et al.ICCV 2025 · 2 citations
- Intern VL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic TasksZhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su et al.CVPR 2024
Related papers
- VideoExplorer: Advancing Long-Horizon Video Understanding via Hierarchical OrchestrationHuaying Yuan, Zheng Liu, Junjie Zhou, Hongjin Qian et al.KDD 2026 · 1 citation
- Linguistic Relative Policy Optimization for Video Anomaly ReasoningJiaxu Leng, Jiankang Zheng, Mengjingcheng Mo, Zhanjie Wu et al.ICML 2026
- Do LVLMs Truly Understand Video Anomalies? Revealing Hallucination via Co-Occurrence PatternsMenghao Zhang, Huazheng Wang, Pengfei Ren, Kangheng Lin et al.NeurIPS 2025 · 4 citations
- Alert-CLIP: Abnormality-aware Latent-Enhanced Representation Tuning of CLIP for Video Anomaly DetectionYiyan Zhu, Menghao Zhang, Haifeng Sun, Pengfei Ren et al.CVPR 2026
- ReAgent-V: A Reward-Driven Multi-Agent Framework for Video UnderstandingYiyang Zhou, Yangfan He, Yaofeng Su, Siwei Han et al.NeurIPS 2025 · 55 citations
