A Distribution Optimization Framework for Confidence Bounds of Risk Measures
Hao Liang, Zhi-Quan Luo
摘要
We present a distribution optimization framework that significantly improves confidence bounds for various risk measures compared to previous methods. Our framework encompasses popular risk measures such as the entropic risk measure, conditional value at risk (CVaR), spectral risk measure, distortion risk measure, equivalent certainty, and rank-dependent expected utility, which are well established in risk-sensitive decision-making literature. To achieve this, we introduce two estimation schemes based on concentration bounds derived from the empirical distribution, specifically using either the Wasserstein distance or the supremum distance. Unlike traditional approaches that add or subtract a confidence radius from the empirical risk measures, our proposed schemes evaluate a specific transformation of the empirical distribution based on the distance. Consequently, our confidence bounds consistently yield tighter results compared to previous methods. We further verify the efficacy of the proposed framework by providing tighter problem-dependent regret bound for the CVaR bandit.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper5
- Being Optimistic to Be Conservative: Quickly Learning a CVaR PolicyRamtin Keramati, Christoph Dann, Alex Tamkin, Emma BrunskillAAAI 2020 · 被引用 86 次
- Universal Off-Policy EvaluationYash Chandak, Scott Niekum, Bruno C. da Silva, Erik G. Learned-Miller 等NeurIPS 2021 · 被引用 64 次
- Concentration bounds for CVaR estimation: The cases of light-tailed and heavy-tailed distributionsPrashanth L. A., Krishna P. Jagannathan, Ravi Kumar KollaICML 2020 · 被引用 53 次
- Optimal Thompson Sampling strategies for support-aware CVaR banditsDorian Baudry, Romain Gautron, Emilie Kaufmann, Odalric MaillardICML 2021 · 被引用 40 次
- Estimation of Spectral Risk MeasuresAjay Kumar Pandey, Prashanth L. A., Sanjay P. BhatAAAI 2021 · 被引用 12 次
相关 Paper
- Regret Bounds for Markov Decision Processes with Recursive Optimized Certainty EquivalentsWenhao Xu, Xuefeng Gao, Xuedong HeICML 2023 · 被引用 14 次
- Optimal Best-Arm Identification Methods for Tail-Risk MeasuresShubhada Agrawal, Wouter M. Koolen, Sandeep JunejaNeurIPS 2021 · 被引用 34 次
- A Unifying Theory of Thompson Sampling for Continuous Risk-Averse BanditsJoel Q. L. Chang, Vincent Y. F. TanAAAI 2022 · 被引用 18 次
- A Reductions Approach to Risk-Sensitive Reinforcement Learning with Optimized Certainty EquivalentsKaiwen Wang, Dawen Liang, Nathan Kallus, Wen SunICML 2025
- Learning Bounds for Risk-sensitive LearningJaeho Lee, Sejun Park, Jinwoo ShinNeurIPS 2020 · 被引用 52 次
