Distributionally-Aware Kernelized Bandit Problems for Risk Aversion
Sho Takemori
摘要
The kernelized bandit problem is a theoretically justified framework and has solid applications to various fields. Recently, there is a growing interest in generalizing the problem to the optimization of risk-averse metrics such as Conditional Value-at-Risk (CVaR) or Mean-Variance (MV). However, due to the model assumption, most existing methods need explicit design of environment random variables and can incur large regret because of possible high dimensionality of them. To address the issues, in this paper, we model the environment using a family of the output distributions (or more precisely, probability kernel) and Kernel Mean Embeddings (KME), and provide novel UCB-type algorithms for CVaR and MV. Moreover, we provide algorithm-independent lower bounds for CVaR in the case of Matérn kernels, and propose a nearly optimal algorithm. Furthermore, we empirically verify our theoretical result in synthetic environments, and demonstrate that our proposed method significantly outperforms a baseline in many cases.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Learning with Good Feature Representations in Bandits and in RL with a Generative ModelTor Lattimore, Csaba Szepesvári, Gellért WeiszICML 2020 · 被引用 181 次
- A Measure-Theoretic Approach to Kernel Conditional Mean EmbeddingsJunhyung Park, Krikamol MuandetNeurIPS 2020 · 被引用 123 次
- Bayesian Optimization of Risk MeasuresSait Cakmak, Raul Astudillo, Peter I. Frazier, Enlu ZhouNeurIPS 2020 · 被引用 65 次
- High-dimensional Experimental Design and Kernel BanditsRomain Camilleri, Kevin Jamieson, Julian Katz-SamuelsICML 2021 · 被引用 63 次
- A Domain-Shrinking based Bayesian Optimization Algorithm with Order-Optimal Regret PerformanceSudeep Salgia, Sattar Vakili, Qing ZhaoNeurIPS 2021 · 被引用 49 次
相关 Paper
- Optimal Best-Arm Identification Methods for Tail-Risk MeasuresShubhada Agrawal, Wouter M. Koolen, Sandeep JunejaNeurIPS 2021 · 被引用 34 次
- Value-at-Risk Optimization with Gaussian ProcessesQuoc Phong Nguyen, Zhongxiang Dai, Bryan Kian Hsiang Low, Patrick JailletICML 2021 · 被引用 34 次
- Optimizing Conditional Value-At-Risk of Black-Box FunctionsQuoc Phong Nguyen, Zhongxiang Dai, Bryan Kian Hsiang Low, Patrick JailletNeurIPS 2021 · 被引用 25 次
- Instance Dependent Regret Analysis of Kernelized BanditsShubhanshu Shekhar, Tara JavidiICML 2022 · 被引用 4 次
- Near-Minimax-Optimal Risk-Sensitive Reinforcement Learning with CVaRKaiwen Wang, Nathan Kallus, Wen SunICML 2023 · 被引用 36 次
