Lune

ICML2026顶会

Generalized Linear Bandits with Memory

Heesang Ann, Hyun-jun Choi, Taehyun Hwang, Younghoon Shin, Haeju Cheong, Min-hwan Oh

2026年份

摘要

We study generalized linear bandits with memory, an endogenous non-stationary setting in which rewards depend on past actions through a finite memory matrix. Building on prior work for linear models Clerici et al.,(2024), we show that the previously known O~(T3/4)\tilde{\mathcal{O}}(T^{3/4}) regret bound stems from a loose analysis, and we provide a sharpened analysis that recovers a O~(T)\tilde{\mathcal{O}}(\sqrt{T}) regret rate in the linear case. We then extend this improvement to generalized linear models and propose a block-wise algorithm based on shrunken confidence bounds. Our algorithm achieves a regret bound of O~(mT+dT+κd2m1/4T1/4+κd2)\tilde{\mathcal{O}}\left(\sqrt{mT} + d\sqrt{T} + \sqrt{\kappa}d^{2} m^{1/4} T^{1/4} + \kappa d^{2} \right), where dd denotes the feature dimension, mm the memory length, and κ\kappa a curvature parameter of the link function. This attains a T\sqrt{T}-type rate despite nonlinear rewards and memory effects. To the best of our knowledge, this analysis provides a unified treatment of memory-induced non-stationarity and nonlinear link functions, while ensuring that the leading regret term is independent of the curvature of the link function. We conduct numerical experiments that are consistent with our theoretical findings.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper11

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖