Open-o3-Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence
Jiahao Meng, Xiangtai Li, Haochen Wang, Tan Yue, Tao Zhang, Lingdong Kong, Yunhai Tong, Anran Wang, Zhiyang Teng, Yujing Wang, Zhuochen Wang
摘要
Most video reasoning models only generate textual reasoning traces without indicating when and where key evidence appears. Recent models such as OpenAI-o3 have sparked wide interest in evidence-centered reasoning for images, yet extending this ability to videos is more challenging due to the need for joint temporal tracking and spatial localization across dynamic scenes. We introduce Open-o3-Video, a non-agent framework that integrates explicit spatio-temporal evidence into video reasoning by highlighting key timestamps, objects, and bounding boxes, making the reasoning process traceable and verifiable. To enable this capability, we first construct high-quality datasets STGR that provide unified spatio-temporal supervision, which is absent in existing resources. We further adopt a cold-start reinforcement learning strategy with specially designed rewards that jointly encourage answer accuracy, temporal alignment, and spatial precision. On the V-STAR benchmark, Open-o3-Video achieves state-of-the-art performance, improving mAM by 14.4% and mLGM by 24.2% over the Qwen2.5-VL baseline, and shows consistent gains across a range of video understanding benchmarks. Beyond accuracy, the grounded reasoning traces produced by Open-o3-Video support confidence-aware test-time scaling, improving answer reliability. The code, model and datasets are publicly available at https://marinero4972.github.io/projects/Open-o3-Video/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- OneThinker: All-in-one Reasoning Model for Image and VideoKaituo Feng, Manyuan Zhang, Hongyu Li, Kaixuan Fan 等CVPR 2026 · 被引用 55 次
- Skyra: AI-Generated Video Detection via Grounded Artifact ReasoningYifei Li, Wenzhao Zheng, Yanran Zhang, Runze Sun 等CVPR 2026 · 被引用 24 次
- REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video UnderstandingJiaze Li, Hao Yin, Wenhui Tan, Jingyang Chen 等CVPR 2026 · 被引用 14 次
- Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop ReasoningXiangyu Zeng, Zhiqiu Zhang, Yuhan Zhu, Xinhao Li 等ICML 2026 · 被引用 14 次
- SAMTok: Representing Any Mask with Two WordsYikang Zhou, Tao Zhang, Dengxian Gong, Yuanzheng Wu 等CVPR 2026 · 被引用 10 次
它引用的顶会 Paper27
- Video-R1: Reinforcing Video Reasoning in MLLMsKaituo Feng, Kaixiong Gong, Bohao Li, Zonghao Guo 等NeurIPS 2025 · 被引用 528 次
- Detecting Moments and Highlights in Videos via Natural Language QueriesJie Lei, Tamara L. Berg, Mohit BansalNeurIPS 2021 · 被引用 425 次
- DeepEyes: Incentivizing "Thinking with Images" via Reinforcement LearningZiwei Zheng, Michael Yang, Jack Hong, Chenxiao Zhao 等ICLR 2026 · 被引用 321 次
- Time-R1: Post-Training Large Vision Language Model for Temporal Video GroundingYe Wang, Ziheng Wang, Boshen Xu, Yang Du 等NeurIPS 2025 · 被引用 143 次
- GRIT: Teaching MLLMs to Think with ImagesYue Fan, Xuehai He, Diji Yang, Kaizhi Zheng 等NeurIPS 2025 · 被引用 132 次
相关 Paper
- VideoTrace-R1: Long Video-based Retrieval-Augmented Generation via Reinforcement LearningZongsheng Cao, Anran Liu, Jun Xie, Feng Chen 等ICML 2026
- MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in VideosKejian Zhu, Zhuoran Jin, Hongbang Yuan, Jiachun Li 等ICLR 2026 · 被引用 22 次
- Conan: Progressive Learning to Reason Like a Detective over Multi-Scale Visual EvidenceKun Ouyang, Yuanxin Liu, Linli Yao, Yishuo Cai 等CVPR 2026 · 被引用 17 次
- VideoSEG-O3: A Multi-turn Reinforcement Learning Framework for Reasoning Video Object SegmentationMing Dai, Sen Yang, Boqiang Duan, Boyuan Tong 等ICML 2026
- ReWatch-R1: Boosting Complex Video Reasoning in Large Vision-Language Models through Agentic Data SynthesisCongzhi Zhang, Zhibin Wang, Yinchao Ma, Jiawei Peng 等ICLR 2026 · 被引用 24 次
