Regime Switching Bandits
Xiang Zhou, Yi Xiong, Ningyuan Chen, Xuefeng Gao
摘要
We study a multi-armed bandit problem where the rewards exhibit regime switching. Specifically, the distributions of the random rewards generated from all arms are modulated by a common underlying state modeled as a finite-state Markov chain. The agent does not observe the underlying state and has to learn the transition matrix and the reward distributions. We propose a learning algorithm for this problem, building on spectral method-of-moments estimations for hidden Markov models, belief error control in partially observable Markov decision processes and upper-confidence-bound methods for online learning. We also establish an upper bound O(T 2/3 √ log T ) for the proposed learning algorithm where T is the learning horizon. Finally, we conduct proof-of-concept experiments to illustrate the performance of the learning algorithm.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Smooth Non-stationary BanditsSu Jia, Qian Xie, Nathan Kallus, Peter I. FrazierICML 2023 · 被引用 14 次
- A Simple and Optimal Policy Design for Online Learning with Safety against Heavy-tailed RiskDavid Simchi-Levi, Zeyu Zheng, Feng ZhuNeurIPS 2022 · 被引用 7 次
- Coordinated Attacks against Contextual Bandits: Fundamental Limits and Defense MechanismsJeongyeol Kwon, Yonathan Efroni, Constantine Caramanis, Shie MannorICML 2022 · 被引用 6 次
- Non-stationary Experimental Design under Linear TrendsDavid Simchi-Levi, Chonghuan Wang, Zeyu ZhengNeurIPS 2023 · 被引用 6 次
- Learning Versatile Skills with Curriculum MaskingYao Tang, Zhihui Xie, Zichuan Lin, Deheng Ye 等NeurIPS 2024 · 被引用 6 次
相关 Paper
- A Direct Approach for Handling Contextual Bandits with Latent State DynamicsZhen Li, Gilles StoltzICML 2026
- Tractable Optimality in Episodic Latent MABsJeongyeol Kwon, Yonathan Efroni, Constantine Caramanis, Shie MannorNeurIPS 2022 · 被引用 3 次
- Online Restless Bandits with Unobserved StatesBowen Jiang, Bo Jiang, Jian Li, Tao Lin 等ICML 2023 · 被引用 8 次
- Adaptive Exploration for Latent-State BanditsJikai Jin, Kenneth Hung, Sanath Kumar Krishnamurthy, Baoyi Shi 等KDD 2026
- Optimistic Whittle Index Policy: Online Learning for Restless BanditsKai Wang, Lily Xu, Aparna Taneja, Milind TambeAAAI 2023 · 被引用 31 次
