MJ-Video: Benchmarking and Rewarding Video Generation with Fine-Grained Video Preference
Haibo Tong, Zhaoyang Wang, Zhaorun Chen, Haonian Ji, Shi Qiu, Siwei Han, Kexin Geng, Zhongkai Xue, Yiyang Zhou, Peng Xia, Mingyu Ding, Rafael Rafailov
Abstract
Recent advancements in video generation have significantly improved the ability to synthesize videos from text instructions. However, existing models still struggle with key challenges such as instruction misalignment, content hallucination, safety concerns, and generation bias. To address these limitations, we introduce MJ-B ENCH -V IDEO , a large-scale video preference benchmark designed to evaluate video generation across five critical aspects: Alignment, Safety, Fineness, Coher-ence & Consistency, and Bias & Fairness . This benchmark further incorporates 28 fine-grained criteria to provide a comprehensive evaluation of video preference. Building upon this dataset, we propose MJ-V IDEO , a Mixture-of-Experts (MoE)- based video reward model designed to deliver fine-grained reward. MJ-V IDEO can dynamically select relevant experts to accurately judge the preference based on the input text-video pair. This architecture enables more precise and adaptable preference judgments. Through extensive benchmarking on MJ-B ENCH -V IDEO , we analyze the limitations of existing video reward models and demonstrate the superior performance of MJ-V IDEO in video preference assessment, achieving 17.58% and 15.87% improvements in overall and fine-grained preference judgments, respectively. Additionally, MJ-V IDEO is able to improve the alignment performance in video generation via preference fine-tuning. Warning: this paper contains content that may be inappropriate or offensive.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e627bce4-5518-4d2a-8abe-6cc9a5bdb011Builds on25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image GenerationYuval Kirstain, Adam Polyak, Uriel Singer, Shahbuland Matiana et al.NeurIPS 2023 · 1,192 citations
- Mixture-of-Experts with Expert Choice RoutingYanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du et al.NeurIPS 2022 · 933 citations
- Training Diffusion Models with Reinforcement LearningKevin Black, Michael Janner, Yilun Du, Ilya Kostrikov et al.ICLR 2024 · 816 citations
- VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric TasksWenhai Wang, Zhe Chen, Xiaokang Chen, Jiannan Wu et al.NeurIPS 2023 · 725 citations
Related papers
- VBench: Comprehensive Benchmark Suite for Video Generative ModelsZiqi Huang, Yinan He, Jiashuo Yu, Fan Zhang et al.CVPR 2024
- VMBench: A Benchmark for Perception-Aligned Video Motion GenerationXinran Ling, Chen Zhu, Meiqi Wu, Hangyu Li et al.ICCV 2025 · 2 citations
- LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text InterpretationJiarui Wang, Huiyu Duan, Ziheng Jia, Zicheng Zhang et al.ICML 2026 · 14 citations
- VideoRealBench: A Chain-of-Thought Realism Evaluation Benchmark for Generated Human-Centric VideosMin Yang, Xinwen Zhang, Jialei Tang, Xin Zhou et al.CVPR 2026
- Generative Ghost: Investigating Ranking Bias Hidden in AI-Generated VideosHaowen Gao, Liang Pang, Shicheng Xu, Leigang Qu et al.ACM MM 2025 · 1 citation
