Wasserstein Actor-Critic: Directed Exploration via Optimism for Continuous-Actions Control
Amarildo Likmeta, Matteo Sacco, Alberto Maria Metelli, Marcello Restelli
Abstract
Uncertainty quantification has been extensively used as a means to achieve efficient directed exploration in Reinforcement Learning (RL). However, state-of-the-art methods for continuous actions still suffer from high sample complexity requirements. Indeed, they either completely lack strategies for propagating the epistemic uncertainty throughout the updates, or they mix it with aleatoric uncertainty while learning the full return distribution (e.g., distributional RL). In this paper, we propose Wasserstein Actor-Critic (WAC), an actor-critic architecture inspired by the recent Wasserstein Q-Learning (WQL), that employs approximate Q-posteriors to represent the epistemic uncertainty and Wasserstein barycenters for uncertainty propagation across the state-action space. WAC enforces exploration in a principled way by guiding the policy learning process with the optimization of an upper bound of the Q-value estimates. Furthermore, we study some peculiar issues that arise when using function approximation, coupled with the uncertainty estimation, and propose a regularized loss for the uncertainty estimation. Finally, we evaluate our algorithm on standard MujoCo tasks as well as suite of continuous-actions domains, where exploration is crucial, in comparison with state-of-the-art baselines. Additional details and results can be found in the supplementary material with our Arxiv preprint.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Estimating Barycenters of Distributions with Neural Optimal TransportAlexander Kolesov, Petr Mokrov, Igor Udovichenko, Milena Gazdieva et al.ICML 2024 · 13 citations
- Wasserstein Policy OptimizationDavid Pfau, Ian Davies, Diana L. Borsa, João Guilherme Madeira Araújo et al.ICML 2025
Builds on7
- SUNRISE: A Simple Unified Framework for Ensemble Learning in Deep Reinforcement LearningKimin Lee, Michael Laskin, Aravind Srinivas, Pieter AbbeelICML 2021 · 239 citations
- Tactical Optimism and Pessimism for Deep Reinforcement LearningTed Moskovitz, Jack Parker-Holder, Aldo Pacchiano, Michael Arbel et al.NeurIPS 2021 · 75 citations
- Task-Agnostic Exploration via Policy Gradient of a Non-Parametric State Entropy EstimateMirco Mutti, Lorenzo Pratissoli, Marcello RestelliAAAI 2021 · 62 citations
- MADE: Exploration via Maximizing Deviation from Explored RegionsTianjun Zhang, Paria Rashidinejad, Jiantao Jiao, Yuandong Tian et al.NeurIPS 2021 · 51 citations
- Adaptive Ensemble Q-learning: Minimizing Estimation Bias via Error FeedbackHang Wang, Sen Lin, Junshan ZhangNeurIPS 2021 · 27 citations
Related papers
- Promoting Stochasticity for Expressive Policies via a Simple and Efficient Regularization MethodQi Zhou, Yufei Kuang, Zherui Qiu, Houqiang Li et al.NeurIPS 2020 · 9 citations
- Bayesian Distributional Policy GradientsLuchen Li, A. Aldo FaisalAAAI 2021 · 11 citations
- Diffusion Actor-Critic with Entropy RegulatorYinuo Wang, Likun Wang, Yuxuan Jiang, Wenjun Zou et al.NeurIPS 2024 · 105 citations
- Taming Aleatoric Impulse in Off-Policy Reinforcement LearningZhouyang Yu, Guojian Zhan, Yang Guan, Jingliang Duan et al.ICML 2026
- Langevin Soft Actor-Critic: Efficient Exploration through Uncertainty-Driven Critic LearningHaque Ishfaq, Guangyuan Wang, Sami Nur Islam, Doina PrecupICLR 2025
