Sharp Analysis for KL-Regularized Contextual Bandits and RLHF
Heyang Zhao, Chenlu Ye, Quanquan Gu, Tong Zhang
Abstract
Reverse-Kullback-Leibler (KL) regularization has emerged to be a predominant technique used to enhance policy optimization in reinforcement learning (RL) and reinforcement learning from human feedback (RLHF), which forces the learned policy to stay close to a reference policy. While the effectiveness and necessity of KL-regularization have been empirically demonstrated in various practical scenarios, current theoretical analysis of KL-regularized RLHF still obtains the same sample complexity as problems without KL-regularization. To understand the fundamental distinction between policy learning objectives with KL-regularization and ones without KL-regularization, we are the first to theoretically demonstrate the power of KL-regularization by providing a sharp analysis for KL-regularized contextual bandits and RLHF, revealing an sample complexity when is sufficiently small. We further explore the role of data coverage in contextual bandits and RLHF. While the coverage assumption is commonly employed in offline RLHF to link the samples from the reference policy to the optimal policy, often at the cost of a multiplicative dependence on the coverage coefficient, its impact on the sample complexity of online RLHF remains unclear. Previous theoretical analyses of online RLHF typically require explicit exploration and additional structural assumptions on the reward function class. In contrast, we show that with sufficient coverage from the reference policy, a simple two-stage mixed sampling strategy can achieve a sample complexity with only an additive dependence on the coverage coefficient. Our results provide a comprehensive understanding of the roles of KL-regularization and data coverage in RLHF, shedding light on the design of more efficient RLHF algorithms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 79f49870-6cb3-46f6-9447-74f749d8f339Cited by top-tier papers10
- Best-of-N through the Smoothing Lens: KL Divergence and Regret AnalysisGholamali Aminian, Idan Shenfeld, Amir R. Asadi, Ahmad Beirami et al.ICLR 2026 · 16 citations
- Greedy Sampling Is Provably Efficient For RLHFDi Wu, Chengshuai Shi, Jing Yang, Cong ShenNeurIPS 2025 · 11 citations
- KL-Regularized RLHF with Multiple Reference Models: Exact Solutions and Sample ComplexityGholamali Aminian, Amir Reza Asadi, Idan Shenfeld, Youssef MrouehNeurIPS 2025 · 11 citations
- Achieving Logarithmic Regret in KL-Regularized Zero-Sum Markov GamesAnupam Nayak, Tong Yang, Osman Yagan, Gauri Joshi et al.ICML 2026 · 9 citations
- Towards a Sharp Analysis of Offline Policy Learning for -Divergence-Regularized Contextual BanditsQingyue Zhao, Kaixuan Ji, Heyang Zhao, Tong Zhang et al.ICLR 2026 · 9 citations
Builds on11
- Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-constraintWei Xiong, Hanze Dong, Chenlu Ye, Ziqi Wang et al.ICML 2024 · 346 citations
- Principled Reinforcement Learning with Human Feedback from Pairwise or K-wise ComparisonsBanghua Zhu, Michael I. Jordan, Jiantao JiaoICML 2023 · 273 citations
- Almost Optimal Model-Free Reinforcement Learningvia Reference-Advantage DecompositionZihan Zhang, Yuan Zhou, Xiangyang JiNeurIPS 2020 · 183 citations
- Online Iterative Reinforcement Learning from Human Feedback with General Preference ModelChenlu Ye, Wei Xiong, Yuheng Zhang, Hanze Dong et al.NeurIPS 2024 · 60 citations
- Corruption-Robust Algorithms with Uncertainty Weighting for Nonlinear Contextual Bandits and Markov Decision ProcessesChenlu Ye, Wei Xiong, Quanquan Gu, Tong ZhangICML 2023 · 40 citations
Related papers
- Logarithmic Regret for Online KL-Regularized Reinforcement LearningHeyang Zhao, Chenlu Ye, Wei Xiong, Quanquan Gu et al.ICML 2025
- Demonstration-Regularized RLDaniil Tiapkin, Denis Belomestny, Daniele Calandriello, Eric Moulines et al.ICLR 2024 · 5 citations
- -Divergence Regularized RLHF: Two Tales of Sampling and Unified AnalysesDi Wu, Chengshuai Shi, Jing Yang, Cong ShenICML 2026
- Can RLHF be More Efficient with Imperfect Reward Models? A Policy Coverage PerspectiveJiawei Huang, Bingcong Li, Christoph Dann, Niao HeICML 2025
- Near-Optimal Regret for KL-Regularized Multi-Armed BanditsKaixuan Ji, Qingyue Zhao, Heyang Zhao, Qiwei Di et al.ICML 2026 · 3 citations
