GMAC: A Distributional Perspective on Actor-Critic Framework
Daniel Wontae Nam, Younghoon Kim, Chan Y. Park
Abstract
In this paper, we devise a distributional framework on actor-critic as a solution to distributional instability, action type restriction, and conflation between samples and statistics. We propose a new method that minimizes the Cramér distance with the multi-step Bellman target distribution generated from a novel Sample-Replacement algorithm denoted SR(), which learns the correct value distribution under multiple Bellman operations. Parameterizing a value distribution with Gaussian Mixture Model further improves the efficiency and the performance of the method, which we name GMAC. We empirically show that GMAC captures the correct representation of value distributions and improves the performance of a conventional actor-critic method with low computational cost, in both discrete and continuous action spaces using Arcade Learning Environment (ALE) and PyBullet environment.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 167a45bd-8b5d-4444-a684-1dd14f3bdef5Cited by top-tier papers4
- Trust Region-Based Safe Distributional Reinforcement Learning for Multiple ConstraintsDohyeong Kim, Kyungjae Lee, Songhwai OhNeurIPS 2023 · 26 citations
- The Nature of Temporal Difference Errors in Multi-step Distributional Reinforcement LearningYunhao Tang, Rémi Munos, Mark Rowland, Bernardo Ávila Pires et al.NeurIPS 2022 · 16 citations
- Quantile Credit AssignmentThomas Mesnard, Wenqi Chen, Alaa Saade, Yunhao Tang et al.ICML 2023 · 3 citations
- Distributional Active InferenceAbdullah Akgül, Gulcin Baykal, Manuel Haussmann, Mustafa Mert Çelikok et al.ICML 2026
Builds on1
Related papers
- CTD4 - a Deep Continuous Distributional Actor-Critic Agent with a Kalman Fusion of Multiple CriticsDavid Valencia, Henry Williams, Yuning Xing, Trevor Gee et al.AAAI 2025 · 6 citations
- Conjugated Discrete Distributions for Distributional Reinforcement LearningBjörn Lindenberg, Jonas Nordqvist, Karl-Olof LindahlAAAI 2022 · 2 citations
- Distributions as Actions: A Unified Framework for Diverse Action SpacesJiamin He, A. Rupam Mahmood, Martha WhiteICLR 2026
- Sample-Efficient Reinforcement Learning via Conservative Model-Based Actor-CriticZhihai Wang, Jie Wang, Qi Zhou, Bin Li et al.AAAI 2022 · 38 citations
- RAMAC: Multimodal Risk-Aware Offline Reinforcement Learning and the Role of Behavior RegularizationKai Fukazawa, Kunal Mundada, Iman SoltaniICML 2026
