A Unifying Theory of Thompson Sampling for Continuous Risk-Averse Bandits
Joel Q. L. Chang, Vincent Y. F. Tan
摘要
This paper unifies the design and the analysis of risk-averse Thompson sampling algorithms for the multi-armed bandit problem for a class of risk functionals ρ that are continuous and dominant. We prove generalised concentration bounds for these continuous and dominant risk functionals and show that a wide class of popular risk functionals belong to this class. Using our newly developed analytical toolkits, we analyse the algorithm ρ-MTS (for multinomial distributions) and prove that they admit asymptotically optimal regret bounds of risk-averse algorithms under the CVaR, proportional hazard, and other ubiquitous risk measures. More generally, we prove the asymptotic optimality of ρ-MTS for Bernoulli distributions for a class of risk measures known as empirical distribution performance measures (EDPMs); this includes the well-known mean-variance. Numerical simulations show that the regret bounds incurred by our algorithms are reasonably tight vis-à-vis algorithm-independent lower bounds.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Stochastic Multi-armed Bandits: Optimal Trade-off among Optimality, Consistency, and Tail RiskDavid Simchi-Levi, Zeyu Zheng, Feng ZhuNeurIPS 2023 · 被引用 8 次
- A Simple and Optimal Policy Design for Online Learning with Safety against Heavy-tailed RiskDavid Simchi-Levi, Zeyu Zheng, Feng ZhuNeurIPS 2022 · 被引用 7 次
- Pricing Experimental Design: Causal Effect, Expected Revenue and Tail RiskDavid Simchi-Levi, Chonghuan WangICML 2023 · 被引用 6 次
- Balancing Risk and Reward: A Batched-Bandit Strategy for Automated Phased ReleaseYufan Li, Jialiang Mao, Iavor BojinovNeurIPS 2023 · 被引用 1 次
- Probably Anytime-Safe Stochastic Combinatorial Semi-BanditsYunlong Hou, Vincent Y. F. Tan, Zixin ZhongICML 2023 · 被引用 1 次
它引用的顶会 Paper4
- Thompson Sampling Algorithms for Mean-Variance BanditsQiuyu Zhu, Vincent Y. F. TanICML 2020 · 被引用 57 次
- Learning Bounds for Risk-sensitive LearningJaeho Lee, Sejun Park, Jinwoo ShinNeurIPS 2020 · 被引用 52 次
- Off-Policy Risk Assessment in Contextual BanditsAudrey Huang, Liu Leqi, Zachary C. Lipton, Kamyar AzizzadenesheliNeurIPS 2021 · 被引用 44 次
- Optimal Thompson Sampling strategies for support-aware CVaR banditsDorian Baudry, Romain Gautron, Emilie Kaufmann, Odalric MaillardICML 2021 · 被引用 40 次
相关 Paper
- Optimal Best-Arm Identification Methods for Tail-Risk MeasuresShubhada Agrawal, Wouter M. Koolen, Sandeep JunejaNeurIPS 2021 · 被引用 34 次
- A Distribution Optimization Framework for Confidence Bounds of Risk MeasuresHao Liang, Zhi-Quan LuoICML 2023 · 被引用 4 次
- Thompson Sampling with Less Exploration is Fast and OptimalTianyuan Jin, Xianglin Yang, Xiaokui Xiao, Pan XuICML 2023 · 被引用 23 次
- Distributionally-Aware Kernelized Bandit Problems for Risk AversionSho TakemoriICML 2022
- Finite-Time Regret of Thompson Sampling Algorithms for Exponential Family Multi-Armed BanditsTianyuan Jin, Pan Xu, Xiaokui Xiao, Anima AnandkumarNeurIPS 2022 · 被引用 19 次
