Boosting Text-to-Video Generative Model with MLLMs Feedback
Xun Wu, Shaohan Huang, Guolong Wang, Jing Xiong, Furu Wei
摘要
Recent advancements in text-to-video generative models, such as Sora [3], have showcased impressive capabilities. These models have attracted significant interest for their potential applications. However, they often rely on extensive datasets of variable quality, which can result in generated videos that lack aesthetic appeal and do not accurately reflect the input text prompts. A promising approach to mitigate these issues is to leverage Reinforcement Learning from Human Feedback (RLHF), which aims to align the outputs of text-to-video models with human preferences. However, the considerable costs associated with manual annotation have led to a scarcity of comprehensive preference datasets. In response to this challenge, our study begins by investigating the efficacy of Multimodal Large Language Models (MLLMs) generated annotations in capturing video preferences, discovering a high degree of concordance with human judgments. Building upon this finding, we utilize MLLMs to perform fine-grained video preference annotations across two dimensions, resulting in the creation of V IDEO P REFER , which includes 135,000 preference annotations. Utilizing this dataset, we introduce V IDEO RM, the first general-purpose reward model tailored for video preference in the text-to-video domain. Our comprehensive experiments confirm the effectiveness of both V IDEO - P REFER and V IDEO RM, representing a significant step forward in the field.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Inference-Time Text-to-Video Alignment with Diffusion Latent Beam SearchYuta Oshima, Masahiro Suzuki, Yutaka Matsuo, Hiroki FurutaNeurIPS 2025 · 被引用 50 次
- FlowMo: Variance-Based Flow Guidance for Coherent Motion in Video GenerationAriel Shaulov, Itay Hazan, Lior Wolf, Hila CheferNeurIPS 2025 · 被引用 22 次
- MultiBanana: A Challenging Benchmark for Multi-Reference Text-to-Image GenerationYuta Oshima, Daiki Miyake, Kohsei Matsutani, Yusuke Iwasawa 等CVPR 2026 · 被引用 10 次
- DreamCS: Geometry-Aware Text-to-3D Generation with Unpaired 3D Reward SupervisionXiandong Zou, Ruihao Xia, Hongsong Wang, Pan ZhouICLR 2026 · 被引用 6 次
- Is Your Video Language Model a Reliable Judge?Ming Liu, Wensheng ZhangICLR 2025
它引用的顶会 Paper28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
相关 Paper
- Tuning Large Multimodal Models for Videos using Reinforcement Learning from AI FeedbackDaechul Ahn, Yura Choi, Youngjae Yu, Dongyeop Kang 等ACL 2024 · 被引用 3 次
- Improving Video Generation with Human FeedbackJie Liu, Gongye Liu, Jiajun Liang, Ziyang Yuan 等NeurIPS 2025 · 被引用 284 次
- LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text InterpretationJiarui Wang, Huiyu Duan, Ziheng Jia, Zicheng Zhang 等ICML 2026 · 被引用 14 次
- VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video GenerationJiazheng Xu, Yu Huang, Jiale Cheng, Yuanming Yang 等AAAI 2026 · 被引用 112 次
- Automated Multi-level Preference for MLLMsMengxi Zhang, Wenhao Wu, Yu Lu, Yuxin Song 等NeurIPS 2024 · 被引用 34 次
