Dual-IPO: Dual-Iterative Preference Optimization for Text-to-Video Generation
Xiaomeng Yang, Mengping Yang, Gong Jia, Luozheng Qin, Zhiyu Tan, Hao Li
Abstract
Recent advances in video generation have enabled thrilling experiences in producing realistic videos driven by scalable diffusion transformers. However, they usually fail to produce satisfactory outputs that are aligned to users' authentic demands and preferences. In this work, we introduce Dual-Iterative Optimization (Dual-IPO), an iterative paradigm that sequentially optimizes both the reward model and the video generation model for improved synthesis quality and human preference alignment. For the reward model, our framework ensures reliable and robust reward signals via CoT-guided reasoning, voting-based self-consistency, and preference certainty estimation. Given this, we optimize video foundation models with guidance of signals from reward model's feedback, thus improving the synthesis quality in subject consistency, motion smoothness and aesthetic quality, etc. The reward model and video generation model complement each other and are progressively improved in the multi-round iteration, without requiring tediously manual preference annotations. Comprehensive experiments demonstrate that the proposed Dual-IPO can effectively and consistently improve the video generation quality of base model with various architectures and sizes, even help a model with only 2B parameters surpass a 5B one. Moreover, our analysis experiments and ablation studies identify the rational of our systematic design and the efficacy of each component.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f7275acf-c2ff-472d-be6f-bdf880db963cCited by top-tier papers4
- Your One-Stop Solution for AI-Generated Video DetectionLong Ma, Zihao Xue, Yan Wang, Zhiyuan Yan et al.CVPR 2026 · 13 citations
- RLGF: Reinforcement Learning with Geometric Feedback for Autonomous Driving Video GenerationTianyi Yan, Wencheng Han, Xia Zhou, Xueyang Zhang et al.NeurIPS 2025 · 9 citations
- Mind the Generative Details: Direct Localized Detail Preference Optimization for Video Diffusion ModelsZitong Huang, Kaidong Zhang, Yukang Ding, Chao Gao et al.CVPR 2026 · 3 citations
- PISCES: Annotation-free Text-to-Video Post-Training via Optimal Transport-Aligned RewardsMinh-Quan Le, Gaurav Mittal, Cheng Zhao, Xianfeng GU et al.ICML 2026 · 2 citations
Builds on32
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
Related papers
- VideoGPA: Distilling Geometry Priors for 3D-Consistent Video GenerationHongyang Du, Hongyang Du, Xiaoyan Cong, Runhao Li et al.ICML 2026 · 16 citations
- InstructVideo: Instructing Video Diffusion Models with Human FeedbackHangjie Yuan, Shiwei Zhang, Xiang Wang, Yujie Wei et al.CVPR 2024 · 13 citations
- Calibrated Multi-Preference Optimization for Aligning Diffusion ModelsKyungmin Lee, Xiahong Li, Qifei Wang, Junfeng He et al.CVPR 2025
- Prompt-A-Video: Prompt your Video Diffusion Model via Preference-Aligned LLMYatai Ji, Jiacheng Zhang, Jie Wu, Shilong Zhang et al.ICCV 2025 · 3 citations
- Identity-Preserving Image-to-Video Generation via Reward-Guided OptimizationLiao Shen, Wentao Jiang, Yiran Zhu, Jiahe Li et al.CVPR 2026 · 8 citations
