Pitfall of Optimism: Distributional Reinforcement Learning by Randomizing Risk Criterion
Taehyun Cho, Seungyub Han, Heesoo Lee, Kyungjae Lee, Jungwoo Lee
摘要
Distributional reinforcement learning algorithms have attempted to utilize estimated uncertainty for exploration, such as optimism in the face of uncertainty. However, using the estimated variance for optimistic exploration may cause biased data collection and hinder convergence or performance. In this paper, we present a novel distributional reinforcement learning algorithm that selects actions by randomizing risk criterion to avoid one-sided tendency on risk. We provide a perturbed distributional Bellman optimality operator by distorting the risk measure and prove the convergence and optimality of the proposed method with the weaker contraction property. Our theoretical results support that the proposed method does not fall into biased exploration and is guaranteed to converge to an optimal return. Finally, we empirically show that our method outperforms other existing distribution-based algorithms in various environments including Atari 55 games.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Distributional Reinforcement Learning with Regularized Wasserstein LossKe Sun, Yingnan Zhao, Wulong Liu, Bei Jiang 等NeurIPS 2024 · 被引用 2 次
- Intrinsic Benefits of Categorical Distributional Loss: Uncertainty-aware Regularized Exploration in Reinforcement LearningKe Sun, Yingnan Zhao, Enze Shi, Yafei Wang 等NeurIPS 2025 · 被引用 1 次
- Distributional Active InferenceAbdullah Akgül, Gulcin Baykal, Manuel Haussmann, Mustafa Mert Çelikok 等ICML 2026
- Bellman Unbiasedness: Toward Provably Efficient Distributional Reinforcement Learning with General Value Function ApproximationTaehyun Cho, Seungyub Han, Seokhun Ju, Dohyeong Kim 等ICML 2025
它引用的顶会 Paper6
- Being Optimistic to Be Conservative: Quickly Learning a CVaR PolicyRamtin Keramati, Christoph Dann, Alex Tamkin, Emma BrunskillAAAI 2020 · 被引用 86 次
- Tactical Optimism and Pessimism for Deep Reinforcement LearningTed Moskovitz, Jack Parker-Holder, Aldo Pacchiano, Michael Arbel 等NeurIPS 2021 · 被引用 75 次
- Risk-Averse Bayes-Adaptive Reinforcement LearningMarc Rigter, Bruno Lacerda, Nick HawesNeurIPS 2021 · 被引用 50 次
- Distributional Reinforcement Learning via Moment MatchingThanh Nguyen-Tang, Sunil Gupta, Svetha VenkateshAAAI 2021 · 被引用 44 次
- Mean-Variance Policy Iteration for Risk-Averse Reinforcement LearningShangtong Zhang, Bo Liu, Shimon WhitesonAAAI 2021 · 被引用 44 次
相关 Paper
- Distributional Reinforcement Learning for Risk-Sensitive PoliciesShiau Hong Lim, Ilyas MalikNeurIPS 2022 · 被引用 54 次
- Variance Control for Distributional Reinforcement LearningQi Kuang, Zhoufan Zhu, Liwen Zhang, Fan ZhouICML 2023 · 被引用 4 次
- ADDQ: Adaptive distributional double Q-learningLeif Döring, Benedikt Wille, Maximilian Birr, Mihail Bîrsan 等ICML 2025
- Stochastically Dominant Distributional Reinforcement LearningJohn D. Martin, Michal Lyskawinski, Xiaohu Li, Brendan J. EnglotICML 2020 · 被引用 26 次
- Regularized Q-learning through Robust AveragingPeter Schmitt-Förster, Tobias SutterICML 2024
