On a Connection Between Imitation Learning and RLHF
Teng Xiao, Yige Yuan, Mingxiao Li, Zhengyu Chen, Vasant G. Honavar
Abstract
This work studies the alignment of large language models with preference data from an imitation learning perspective. We establish a close theoretical connection between reinforcement learning from human feedback (RLHF) and imitation learning (IL), revealing that RLHF implicitly performs imitation learning on the preference data distribution. Building on this connection, we propose DIL, a principled framework that directly optimizes the imitation learning objective. DIL provides a unified imitation learning perspective on alignment, encompassing existing alignment algorithms as special cases while naturally introducing new variants. By bridging IL and RLHF, DIL offers new insights into alignment with RLHF. Extensive experiments demonstrate that DIL outperforms existing methods on various challenging benchmarks. The code for DIL is available at https://github.com/tengxiao1/DIL .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9a3c8fe9-f175-4106-b56d-278de8d5f290Cited by top-tier papers12
- All Roads Lead to Likelihood: The Value of Reinforcement Learning in Fine-TuningGokul Swamy, Sanjiban Choudhury, Wen Sun, Steven Wu et al.ICLR 2026 · 66 citations
- Doubly Robust Alignment for Large Language ModelsErhan Xu, Kai Ye, Hongyi Zhou, Luhan Zhu et al.NeurIPS 2025 · 14 citations
- Inference-time Alignment in Continuous SpaceYige Yuan, Teng Xiao, Yunfan Li, Bingbing Xu et al.NeurIPS 2025 · 9 citations
- Simple Distillation for One-Step Diffusion ModelsHuaisheng Zhu, Teng Xiao, Shijie Zhou, Zhimeng Guo et al.NeurIPS 2025 · 7 citations
- From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time AlignmentBin Xie, Bingbing Xu, Yige Yuan, Shengmao Zhu et al.ACL 2025 · 4 citations
Builds on20
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- Model Alignment as Prospect Theoretic OptimizationKawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky et al.ICML 2024 · 973 citations
- Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine TranslationHaoran Xu, Amr Sharaf, Yunmo Chen, Weiting Tan et al.ICML 2024 · 447 citations
Related papers
- Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-constraintWei Xiong, Hanze Dong, Chenlu Ye, Ziqi Wang et al.ICML 2024 · 346 citations
- Improving LLM General Preference Alignment via Optimistic Online Mirror DescentYuheng Zhang, Dian Yu, Tao Ge, Linfeng Song et al.NeurIPS 2025 · 27 citations
- Approximated Variational Bayesian Inverse Reinforcement Learning for Large Language Model AlignmentYuang Cai, Yuyu Yuan, Jinsheng Shi, Qinhong LinAAAI 2025 · 5 citations
- RLCD: Reinforcement Learning from Contrastive Distillation for LM AlignmentKevin Yang, Dan Klein, Asli Celikyilmaz, Nanyun Peng et al.ICLR 2024 · 37 citations
- Influence-based Online Experience Selection for Effective RLHFYifan Gong, Jing Yao, Xiting Wang, Xunlong Wang et al.ACL 2026
