Stochastic Multi-Armed Bandits with Control Variates
Arun Verma, Manjesh Kumar Hanawal
摘要
This paper studies a new variant of the stochastic multi-armed bandits problem where auxiliary information about the arm rewards is available in the form of control variates. In many applications like queuing and wireless networks, the arm rewards are functions of some exogenous variables. The mean values of these variables are known a priori from historical data and can be used as control variates. Leveraging the theory of control variates, we obtain mean estimates with smaller variance and tighter confidence bounds. We develop an upper confidence bound based algorithm named UCB-CV and characterize the regret bounds in terms of the correlation between rewards and control variates when they follow a multivariate normal distribution. We also extend UCB-CV to other distributions using resampling methods like Jackknifing and Splitting. Experiments on synthetic problem instances validate performance guarantees of the proposed algorithms.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Exploiting Correlated Auxiliary Feedback in Parameterized BanditsArun Verma, Zhongxiang Dai, Yao Shu, Bryan Kian Hsiang LowNeurIPS 2023 · 被引用 6 次
- Accelerating Unbiased LLM Evaluation via Synthetic FeedbackZhaoyi Zhou, Yuda Song, Andrea ZanetteICML 2025
相关 Paper
- Precise Asymptotics and Refined Regret of Variance-Aware UCBYingying Fan, Yuxuan Han, Jinchi Lv, Xiaocong Xu 等NeurIPS 2025 · 被引用 5 次
- Uplifting BanditsYu-Guan Hsieh, Shiva Prasad Kasiviswanathan, Branislav KvetonNeurIPS 2022 · 被引用 2 次
- Distributionally-Aware Kernelized Bandit Problems for Risk AversionSho TakemoriICML 2022
- Mixed-Effects Contextual BanditsKyungbok Lee, Myunghee Cho Paik, Min-hwan Oh, Gi-Soo KimAAAI 2024 · 被引用 2 次
- Robust Contextual Bandits via BootstrappingQiao Tang, Hong Xie, Yunni Xia, Jia Lee 等AAAI 2021 · 被引用 2 次
