Thinking with Frames: Generative Video Distortion Evaluation via Frame Reward Model
Yuan Wang, Borui Liao, Huijuan Huang, Jinda Lu, Ouxiang Li, Kuien Liu, Meng Wang, Xiang Wang
Abstract
Recent advances in video reward models and post-training strategies have improved text-to-video (T2V) generation. While these models typically assess visual quality, motion quality, and text alignment, they often overlook key structural distortions, such as abnormal object appearances and interactions, which can degrade the overall quality of the generative video. To address this gap, we introduce REACT, a frame-level reward model designed specifically for structural distortions evaluation in generative videos. REACT assigns point-wise scores and attribution labels by reasoning over video frames, focusing on recognizing distortions. To support this, we construct a large-scale human preference dataset, annotated based on our proposed taxonomy of structural distortions, and generate additional data using a efficient Chain-of-Thought (CoT) synthesis pipeline. REACT is trained with a two-stage framework: (1) supervised fine-tuning with masked loss for domain knowledge injection, followed by (2) reinforcement learning with Group Relative Policy Optimization (GRPO) and pairwise rewards to enhance reasoning capability and align output scores with human preferences. During inference, a dynamic sampling mechanism is introduced to focus on frames most likely to exhibit distortion. We also present REACT-Bench, a benchmark for generative video distortion evaluation. Experimental results demonstrate that REACT complements existing reward models in assessing structutal distortion, achieving both accurate quantitative evaluations and interpretable attribution analysis.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext af0396f5-098b-4e7a-865b-1d918febce8dCited by top-tier papers2
- Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?Ouxiang Li, Yuan Wang, Xinting Hu, Huijuan Huang et al.ICLR 2026 · 39 citations
- SPEED: Scalable, Precise, and Efficient Concept Erasure for Diffusion ModelsOuxiang Li, Yuan Wang, Xinting Hu, Houcheng Jiang et al.ICLR 2026 · 37 citations
Builds on27
- Flow-GRPO: Training Flow Matching Models via Online RLJie Liu, Gongye Liu, Jiajun Liang, Yangguang Li et al.NeurIPS 2025 · 647 citations
- Visual-RFT: Visual Reinforcement Fine-TuningZiyu Liu, Zeyi Sun, Yuhang Zang, Xiaoyi Dong et al.ICCV 2025 · 563 citations
- Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined LevelsHaoning Wu, Zicheng Zhang, Weixia Zhang, Chaofeng Chen et al.ICML 2024 · 499 citations
- Improving Video Generation with Human FeedbackJie Liu, Gongye Liu, Jiajun Liang, Ziyang Yuan et al.NeurIPS 2025 · 284 citations
- Improve Vision Language Model Chain-of-thought ReasoningRuohong Zhang, Bowen Zhang, Yanghao Li, Haotian Zhang et al.ACL 2025 · 135 citations
Related papers
- VQAThinker: Exploring Generalizable and Explainable Video Quality Assessment via Reinforcement LearningLinhan Cao, Wei Sun, Weixia Zhang, Xiangyang Zhu et al.AAAI 2026 · 6 citations
- ReWatch-R1: Boosting Complex Video Reasoning in Large Vision-Language Models through Agentic Data SynthesisCongzhi Zhang, Zhibin Wang, Yinchao Ma, Jiawei Peng et al.ICLR 2026 · 24 citations
- SoliReward: Mitigating Susceptibility to Reward Hacking and Annotation Noise in Video Generation Reward ModelsJiesong Lian, Ruizhe Zhong, Zixiang Zhou, Xiaoyue Mi et al.CVPR 2026 · 3 citations
- MVP: Enhancing Video Large Language Models via Self-supervised Masked Video PredictionXiaokun Sun, Zezhong Wu, Zewen Ding, Linli XuACL 2026 · 1 citation
- Seeing What Matters: Visual Preference Policy Optimization for Visual GenerationZiqi Ni, Yuanzhi Liang, Rui Li, Yi Zhou et al.CVPR 2026 · 9 citations
