Tuning Large Multimodal Models for Videos using Reinforcement Learning from AI Feedback
Daechul Ahn, Yura Choi, Youngjae Yu, Dongyeop Kang, Jonghyun Choi
Abstract
Recent advancements in large language models have influenced the development of video large multimodal models (VLMMs). Previous approaches for VLMMs involve Supervised Fine-Tuning (SFT) with instruction-tuned datasets, integrating LLM with visual encoders, and additional learnable parameters. Here, aligning video with text, and vice versa, remains a challenge, primarily due to the insufficient quality and quantity of multimodal instructiontune data compared to that of text-only. This discrepancy often results in alignments that poorly ground the video content. To address this, we present a novel alignment strategy that employs a multimodal AI system equipped with Reinforcement Learning from AI Feedback (RLAIF), providing self-preference feedback to refine itself and facilitating the alignment of video and text modalities. Our approach uniquely integrates detailed video descriptions as context into a multimodal AI system during preference feedback generation to enrich the understanding of video content, a process we call context-aware reward modeling. Empirical evaluations on various video benchmarks demonstrate that our VLM-RLAIF outperforms existing approaches, including the SFT model. We commit to open-sourcing our code, models, and datasets to foster further research in this area. https://github.com/ yonseivnl/vlm-rlaif
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1752e0e8-ec0a-4adb-8cc3-9b0cc763c2f8Cited by top-tier papers13
- MMDuet2: Enhancing Proactive Interaction of Video MLLMs with Multi-Turn Reinforcement LearningYueqian Wang, Songxiang Liu, Disong Wang, Nuo Xu et al.ICLR 2026 · 21 citations
- Flow4Agent: Long-form Video Understanding via Motion Prior from Optical FlowRuyang Liu, Shangkun Sun, Haoran Tang, Wei Gao et al.ICCV 2025 · 15 citations
- VidChain: Chain-of-Tasks with Metric-based Direct Preference Optimization for Dense Video CaptioningJi Soo Lee, Jongha Kim, Jeehye Na, Jinyoung Park et al.AAAI 2025 · 11 citations
- ISR-DPO: Aligning Large Multimodal Models for Videos by Iterative Self-Retrospective DPODaechul Ahn, Yura Choi, San Kim, Youngjae Yu et al.AAAI 2025 · 5 citations
- VideoPASTA: 7K Preference Pairs That Matter for Video-LLM AlignmentYogesh Kulkarni, Pooyan FazliEMNLP 2025 · 5 citations
Builds on17
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-trainingLinjie Li, Yen-Chun Chen, Yu Cheng, Zhe Gan et al.EMNLP 2020 · 387 citations
Related papers
- Boosting Text-to-Video Generative Model with MLLMs FeedbackXun Wu, Shaohan Huang, Guolong Wang, Jing Xiong et al.NeurIPS 2024 · 23 citations
- Re-Align: Aligning Vision Language Models via Retrieval-Augmented Direct Preference OptimizationShuo Xing, Peiran Li, Yuping Wang, Ruizheng Bai et al.EMNLP 2025 · 2 citations
- VQ-Insight: Teaching VLMs for AI-Generated Video Quality Understanding via Progressive Visual Reinforcement LearningXuanyu Zhang, Weiqi Li, Shijie Zhao, Junlin Li et al.AAAI 2026 · 20 citations
- Approximated Variational Bayesian Inverse Reinforcement Learning for Large Language Model AlignmentYuang Cai, Yuyu Yuan, Jinsheng Shi, Qinhong LinAAAI 2025 · 5 citations
- LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text InterpretationJiarui Wang, Huiyu Duan, Ziheng Jia, Zicheng Zhang et al.ICML 2026 · 14 citations
