Lune

ICML2025顶会

Logarithmic Regret for Online KL-Regularized Reinforcement Learning

Heyang Zhao, Chenlu Ye, Wei Xiong, Quanquan Gu, Tong Zhang

出版方
2025年份
9顶会引用

摘要

Recent advances in Reinforcement Learning from Human Feedback (RLHF) have shown that KLregularization plays a pivotal role in improving the efficiency of RL fine-tuning for large language models (LLMs). Despite its empirical advantage, the theoretical difference between KL-regularized RL and standard RL remains largely under-explored. While there is a recent line of work on the theoretical analysis of KLregularized objective in decision making (Xiong et al., 2024a; Xie et al., 2024; Zhao et al., 2024) , these analyses either reduce to the traditional RL setting or rely on strong coverage assumptions. In this paper, we propose an optimismbased KL-regularized online contextual bandit algorithm, and provide a novel analysis of its regret. By carefully leveraging the benign optimization landscape induced by the KL-regularization and the optimistic reward estimation, our algorithm achieves an O η log(N R T ) • d R logarithmic regret bound, where η, N R , T, d R denote the KLregularization parameter, the cardinality of the reward function class, number of rounds, and the complexity of the reward function class. Furthermore, we extend our algorithm and analysis to reinforcement learning by developing a novel decomposition over transition steps and also obtain a similar logarithmic regret bound.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper9

问问它们各自怎么用它

它引用的顶会 Paper26

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖