Sharp Analysis for KL-Regularized Contextual Bandits and RLHF
Heyang Zhao, Chenlu Ye, Quanquan Gu, Tong Zhang
摘要
Reverse-Kullback-Leibler (KL) regularization has emerged to be a predominant technique used to enhance policy optimization in reinforcement learning (RL) and reinforcement learning from human feedback (RLHF), which forces the learned policy to stay close to a reference policy. While the effectiveness and necessity of KL-regularization have been empirically demonstrated in various practical scenarios, current theoretical analysis of KL-regularized RLHF still obtains the same sample complexity as problems without KL-regularization. To understand the fundamental distinction between policy learning objectives with KL-regularization and ones without KL-regularization, we are the first to theoretically demonstrate the power of KL-regularization by providing a sharp analysis for KL-regularized contextual bandits and RLHF, revealing an sample complexity when is sufficiently small. We further explore the role of data coverage in contextual bandits and RLHF. While the coverage assumption is commonly employed in offline RLHF to link the samples from the reference policy to the optimal policy, often at the cost of a multiplicative dependence on the coverage coefficient, its impact on the sample complexity of online RLHF remains unclear. Previous theoretical analyses of online RLHF typically require explicit exploration and additional structural assumptions on the reward function class. In contrast, we show that with sufficient coverage from the reference policy, a simple two-stage mixed sampling strategy can achieve a sample complexity with only an additive dependence on the coverage coefficient. Our results provide a comprehensive understanding of the roles of KL-regularization and data coverage in RLHF, shedding light on the design of more efficient RLHF algorithms.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Best-of-N through the Smoothing Lens: KL Divergence and Regret AnalysisGholamali Aminian, Idan Shenfeld, Amir R. Asadi, Ahmad Beirami 等ICLR 2026 · 被引用 16 次
- Greedy Sampling Is Provably Efficient For RLHFDi Wu, Chengshuai Shi, Jing Yang, Cong ShenNeurIPS 2025 · 被引用 11 次
- KL-Regularized RLHF with Multiple Reference Models: Exact Solutions and Sample ComplexityGholamali Aminian, Amir Reza Asadi, Idan Shenfeld, Youssef MrouehNeurIPS 2025 · 被引用 11 次
- Achieving Logarithmic Regret in KL-Regularized Zero-Sum Markov GamesAnupam Nayak, Tong Yang, Osman Yagan, Gauri Joshi 等ICML 2026 · 被引用 9 次
- Towards a Sharp Analysis of Offline Policy Learning for -Divergence-Regularized Contextual BanditsQingyue Zhao, Kaixuan Ji, Heyang Zhao, Tong Zhang 等ICLR 2026 · 被引用 9 次
它引用的顶会 Paper11
- Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-constraintWei Xiong, Hanze Dong, Chenlu Ye, Ziqi Wang 等ICML 2024 · 被引用 346 次
- Principled Reinforcement Learning with Human Feedback from Pairwise or K-wise ComparisonsBanghua Zhu, Michael I. Jordan, Jiantao JiaoICML 2023 · 被引用 273 次
- Almost Optimal Model-Free Reinforcement Learningvia Reference-Advantage DecompositionZihan Zhang, Yuan Zhou, Xiangyang JiNeurIPS 2020 · 被引用 183 次
- Online Iterative Reinforcement Learning from Human Feedback with General Preference ModelChenlu Ye, Wei Xiong, Yuheng Zhang, Hanze Dong 等NeurIPS 2024 · 被引用 60 次
- Corruption-Robust Algorithms with Uncertainty Weighting for Nonlinear Contextual Bandits and Markov Decision ProcessesChenlu Ye, Wei Xiong, Quanquan Gu, Tong ZhangICML 2023 · 被引用 40 次
相关 Paper
- Logarithmic Regret for Online KL-Regularized Reinforcement LearningHeyang Zhao, Chenlu Ye, Wei Xiong, Quanquan Gu 等ICML 2025
- Demonstration-Regularized RLDaniil Tiapkin, Denis Belomestny, Daniele Calandriello, Eric Moulines 等ICLR 2024 · 被引用 5 次
- -Divergence Regularized RLHF: Two Tales of Sampling and Unified AnalysesDi Wu, Chengshuai Shi, Jing Yang, Cong ShenICML 2026
- Can RLHF be More Efficient with Imperfect Reward Models? A Policy Coverage PerspectiveJiawei Huang, Bingcong Li, Christoph Dann, Niao HeICML 2025
- Near-Optimal Regret for KL-Regularized Multi-Armed BanditsKaixuan Ji, Qingyue Zhao, Heyang Zhao, Qiwei Di 等ICML 2026 · 被引用 3 次
