Unified Reinforcement and Imitation Learning for Vision-Language Models
Byung-Kwan Lee, Ryo Hachiuma, Yong Man Ro, Yu-Chiang Frank Wang, Yueh-Hua Wu
Abstract
Vision-Language Models (VLMs) have achieved remarkable progress, yet their large scale often renders them impractical for resource-constrained environments. This paper introduces Unified Reinforcement and Imitation Learning (RIL), a novel and efficient training algorithm designed to create powerful, lightweight VLMs. RIL distinctively combines the strengths of reinforcement learning with adversarial imitation learning. This enables smaller student VLMs not only to mimic the sophisticated text generation of large teacher models but also to systematically improve their generative capabilities through reinforcement signals. Key to our imitation framework is an LLM-based discriminator that adeptly distinguishes between student and teacher outputs, complemented by guidance from multiple large teacher VLMs to ensure diverse learning. This unified learning strategy, leveraging both reinforcement and imitation, empowers student models to achieve significant performance gains, making them competitive with leading closed-source VLMs. Extensive experiments on diverse vision-language benchmarks demonstrate that RIL significantly narrows the performance gap with state-of-the-art open-and closed-source VLMs and, in several instances, surpasses them. [Project Page] However, the practical implementation of RIL for VLMs entails several specific challenges. Firstly, relying on continuous discriminator scores (ranging from zero to one) to assess similarity to large VLM outputs can introduce ambiguity into the learning signal. To ensure a clearer and more decisive reward, and drawing inspiration from prior works that binarize answer rewards [7, 43] , we convert the discriminator's similarity score into a binary value. Secondly, the discriminator, by design, focuses on stylistic similarity and does not inherently verify factual correctness against ground truth
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d91cf63e-fd29-494c-a99e-c10ed93ebcb4Cited by top-tier papers2
- Fast-ThinkAct: Efficient Vision-Language-Action Reasoning via Verbalizable Latent PlanningChi-Pin Huang, Yunze Man, Zhiding Yu, Min-Hung Chen et al.CVPR 2026 · 24 citations
- Masking Teacher and Reinforcing Student for Distilling Vision-Language ModelsByung-Kwan Lee, Yu-Chiang Frank Wang, Ryo HachiumaCVPR 2026 · 7 citations
Builds on41
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu et al.NeurIPS 2022 · 2,727 citations
Related papers
- Unified Generative and Discriminative Training for Multi-modal Large Language ModelsWei Chow, Juncheng Li, Qifan Yu, Kaihang Pan et al.NeurIPS 2024 · 19 citations
- VladVA: Discriminative Fine-tuning of LVLMsYassine Ouali, Adrian Bulat, Alexandros Xenos, Anestis Zaganidis et al.CVPR 2025
- TextGAIL: Generative Adversarial Imitation Learning for Text GenerationQingyang Wu, Lei Li, Zhou YuAAAI 2021 · 54 citations
- See First, Reason Later: Mutual Information-Guided Reinforcement Learning for Vision-Language ModelsYin Zhang, Zonghan Wu, Jiaxuan Zhao, Junfeng Fang et al.ICML 2026
- ILLUME: Illuminating Your LLMs to See, Draw, and Self-EnhanceChunwei Wang, Guansong Lu, Junwei Yang, Runhui Huang et al.ICCV 2025 · 5 citations
