Wasserstein Actor-Critic: Directed Exploration via Optimism for Continuous-Actions Control
Amarildo Likmeta, Matteo Sacco, Alberto Maria Metelli, Marcello Restelli
摘要
Uncertainty quantification has been extensively used as a means to achieve efficient directed exploration in Reinforcement Learning (RL). However, state-of-the-art methods for continuous actions still suffer from high sample complexity requirements. Indeed, they either completely lack strategies for propagating the epistemic uncertainty throughout the updates, or they mix it with aleatoric uncertainty while learning the full return distribution (e.g., distributional RL). In this paper, we propose Wasserstein Actor-Critic (WAC), an actor-critic architecture inspired by the recent Wasserstein Q-Learning (WQL), that employs approximate Q-posteriors to represent the epistemic uncertainty and Wasserstein barycenters for uncertainty propagation across the state-action space. WAC enforces exploration in a principled way by guiding the policy learning process with the optimization of an upper bound of the Q-value estimates. Furthermore, we study some peculiar issues that arise when using function approximation, coupled with the uncertainty estimation, and propose a regularized loss for the uncertainty estimation. Finally, we evaluate our algorithm on standard MujoCo tasks as well as suite of continuous-actions domains, where exploration is crucial, in comparison with state-of-the-art baselines. Additional details and results can be found in the supplementary material with our Arxiv preprint.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Estimating Barycenters of Distributions with Neural Optimal TransportAlexander Kolesov, Petr Mokrov, Igor Udovichenko, Milena Gazdieva 等ICML 2024 · 被引用 13 次
- Wasserstein Policy OptimizationDavid Pfau, Ian Davies, Diana L. Borsa, João Guilherme Madeira Araújo 等ICML 2025
它引用的顶会 Paper7
- SUNRISE: A Simple Unified Framework for Ensemble Learning in Deep Reinforcement LearningKimin Lee, Michael Laskin, Aravind Srinivas, Pieter AbbeelICML 2021 · 被引用 239 次
- Tactical Optimism and Pessimism for Deep Reinforcement LearningTed Moskovitz, Jack Parker-Holder, Aldo Pacchiano, Michael Arbel 等NeurIPS 2021 · 被引用 75 次
- Task-Agnostic Exploration via Policy Gradient of a Non-Parametric State Entropy EstimateMirco Mutti, Lorenzo Pratissoli, Marcello RestelliAAAI 2021 · 被引用 62 次
- MADE: Exploration via Maximizing Deviation from Explored RegionsTianjun Zhang, Paria Rashidinejad, Jiantao Jiao, Yuandong Tian 等NeurIPS 2021 · 被引用 51 次
- Adaptive Ensemble Q-learning: Minimizing Estimation Bias via Error FeedbackHang Wang, Sen Lin, Junshan ZhangNeurIPS 2021 · 被引用 27 次
相关 Paper
- Promoting Stochasticity for Expressive Policies via a Simple and Efficient Regularization MethodQi Zhou, Yufei Kuang, Zherui Qiu, Houqiang Li 等NeurIPS 2020 · 被引用 9 次
- Bayesian Distributional Policy GradientsLuchen Li, A. Aldo FaisalAAAI 2021 · 被引用 11 次
- Diffusion Actor-Critic with Entropy RegulatorYinuo Wang, Likun Wang, Yuxuan Jiang, Wenjun Zou 等NeurIPS 2024 · 被引用 105 次
- Taming Aleatoric Impulse in Off-Policy Reinforcement LearningZhouyang Yu, Guojian Zhan, Yang Guan, Jingliang Duan 等ICML 2026
- Langevin Soft Actor-Critic: Efficient Exploration through Uncertainty-Driven Critic LearningHaque Ishfaq, Guangyuan Wang, Sami Nur Islam, Doina PrecupICLR 2025
