Dual-IPO: Dual-Iterative Preference Optimization for Text-to-Video Generation
Xiaomeng Yang, Mengping Yang, Gong Jia, Luozheng Qin, Zhiyu Tan, Hao Li
摘要
Recent advances in video generation have enabled thrilling experiences in producing realistic videos driven by scalable diffusion transformers. However, they usually fail to produce satisfactory outputs that are aligned to users' authentic demands and preferences. In this work, we introduce Dual-Iterative Optimization (Dual-IPO), an iterative paradigm that sequentially optimizes both the reward model and the video generation model for improved synthesis quality and human preference alignment. For the reward model, our framework ensures reliable and robust reward signals via CoT-guided reasoning, voting-based self-consistency, and preference certainty estimation. Given this, we optimize video foundation models with guidance of signals from reward model's feedback, thus improving the synthesis quality in subject consistency, motion smoothness and aesthetic quality, etc. The reward model and video generation model complement each other and are progressively improved in the multi-round iteration, without requiring tediously manual preference annotations. Comprehensive experiments demonstrate that the proposed Dual-IPO can effectively and consistently improve the video generation quality of base model with various architectures and sizes, even help a model with only 2B parameters surpass a 5B one. Moreover, our analysis experiments and ablation studies identify the rational of our systematic design and the efficacy of each component.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Your One-Stop Solution for AI-Generated Video DetectionLong Ma, Zihao Xue, Yan Wang, Zhiyuan Yan 等CVPR 2026 · 被引用 13 次
- RLGF: Reinforcement Learning with Geometric Feedback for Autonomous Driving Video GenerationTianyi Yan, Wencheng Han, Xia Zhou, Xueyang Zhang 等NeurIPS 2025 · 被引用 9 次
- Mind the Generative Details: Direct Localized Detail Preference Optimization for Video Diffusion ModelsZitong Huang, Kaidong Zhang, Yukang Ding, Chao Gao 等CVPR 2026 · 被引用 3 次
- PISCES: Annotation-free Text-to-Video Post-Training via Optimal Transport-Aligned RewardsMinh-Quan Le, Gaurav Mittal, Cheng Zhao, Xianfeng GU 等ICML 2026 · 被引用 2 次
它引用的顶会 Paper32
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
相关 Paper
- VideoGPA: Distilling Geometry Priors for 3D-Consistent Video GenerationHongyang Du, Hongyang Du, Xiaoyan Cong, Runhao Li 等ICML 2026 · 被引用 16 次
- InstructVideo: Instructing Video Diffusion Models with Human FeedbackHangjie Yuan, Shiwei Zhang, Xiang Wang, Yujie Wei 等CVPR 2024 · 被引用 13 次
- Calibrated Multi-Preference Optimization for Aligning Diffusion ModelsKyungmin Lee, Xiahong Li, Qifei Wang, Junfeng He 等CVPR 2025
- Prompt-A-Video: Prompt your Video Diffusion Model via Preference-Aligned LLMYatai Ji, Jiacheng Zhang, Jie Wu, Shilong Zhang 等ICCV 2025 · 被引用 3 次
- Identity-Preserving Image-to-Video Generation via Reward-Guided OptimizationLiao Shen, Wentao Jiang, Yiran Zhu, Jiahe Li 等CVPR 2026 · 被引用 8 次
