Lune

ICLR2025顶会

Second Order Bounds for Contextual Bandits with Function Approximation

Aldo Pacchiano

2025年份
3顶会引用

摘要

Many works have developed no-regret algorithms for contextual bandits with function approximation, where the mean rewards over context-action pairs belong to a function class F. Although there are many approaches to this problem, algorithms based on the principle of optimism, such as optimistic least squares have gained in importance. The regret of optimistic least squares scales as r O ´ad eluder pFq logpFqT where d eluder pFq is a statistical measure of the complexity of the function class F known as eluder dimension. Unfortunately, even if the variance of the measurement noise of the rewards at time t equals σ 2 t and these are close to zero, the optimistic least squares algorithm's regret scales with ? T . In this work we are the first to develop algorithms that satisfy regret bounds for contextual bandits with function approximation of the form r O ´σa logpFqd eluder pFqT deluder pFq ¨logp|F|q ¯when the variances are unknown and satisfy σ 2 t " σ for all t and r O ˆdeluder pFq b logpFq ř T t"1 σ 2 t deluder pFq ¨logp|F|q ẇhen the variances change at every time-step. These bounds generalize existing techniques for deriving second order bounds in contextual linear problems.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper3

问问它们各自怎么用它

它引用的顶会 Paper8

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖