Beyond the Last Frame: Process-aware Evaluation for Generative Video Reasoning
Yifan Li, Yukai Gu, Yingqian Min, Zikang Liu, Yifan Du, Kun Zhou, Min Yang, Xin Zhao, Minghui Qiu
Abstract
Recent breakthroughs in video generation have demonstrated an emerging capability termed Chain-of-Frames (CoF) reasoning, where models resolve complex tasks through the generation of continuous frames. While these models show promise for Generative Video Reasoning (GVR), existing evaluation frameworks often rely on single-frame assessments, which can lead to outcome-hacking, where a model reaches a correct conclusion through an erroneous process. To address this, we propose a process-aware evaluation paradigm. We introduce VIPER, a comprehensive benchmark spanning 16 tasks across temporal, structural, symbolic, spatial, physics, and planning reasoning. Furthermore, we propose Process-outcome Consistency (POC@r), a new metric that utilizes VLM-as-Judge with a hierarchical rubric to evaluate both the validity of the intermediate steps and the final result. Our experiments reveal that state-of-the-art video models achieve POC@1.0 only about 20% and exhibit a significant outcome-hacking. We further explore the impact of test-time scaling and sampling robustness, highlighting a substantial gap between current video generation and true generalized visual reasoning. Our benchmark are released at https://github.com/RUCAIBox/VIPER .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on7
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual ContextsPan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu et al.ICLR 2024 · 1,472 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
- MMMU: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert AGIXiang Yue, Yuansheng Ni, Tianyu Zheng, Kai Zhang et al.CVPR 2024 · 213 citations
- Pathways on the Image Manifold: Image Editing via Video GenerationNoam Rotstein, Gal Yona, Daniel Silver, Roy Velich et al.CVPR 2025
- Pyramidal Flow Matching for Efficient Video Generative ModelingYang Jin, Zhicheng Sun, Ningyuan Li, Kun Xu et al.ICLR 2025
Related papers
- TiViBench: Benchmarking Think-in-Video Reasoning for Video GenerationHarold Haodong Chen, Disen Lan, Wen-Jie Shu, Qingyang Liu et al.CVPR 2026
- VideoReasonBench: Can MLLMs Perform Vision-Centric Complex Video Reasoning?Yuanxin Liu, Kun Ouyang, Haoning Wu, Yi Liu et al.ICLR 2026 · 20 citations
- A Very Big Video Reasoning SuiteMaijunxian Wang, Ruisi Wang, Juyi Lin, Ran Ji et al.ICML 2026 · 20 citations
- VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative VideosJiashuo Yu, Yue Wu, Meng Chu, Zhifei Ren et al.ICCV 2025 · 3 citations
- SVBench: Evaluation of Video Generation Models on Social ReasoningWenshuo Peng, Gongxuan Wang, Tianmeng Yang, Chuanhao Li et al.CVPR 2026 · 5 citations
