GMAC: A Distributional Perspective on Actor-Critic Framework
Daniel Wontae Nam, Younghoon Kim, Chan Y. Park
摘要
In this paper, we devise a distributional framework on actor-critic as a solution to distributional instability, action type restriction, and conflation between samples and statistics. We propose a new method that minimizes the Cramér distance with the multi-step Bellman target distribution generated from a novel Sample-Replacement algorithm denoted SR(), which learns the correct value distribution under multiple Bellman operations. Parameterizing a value distribution with Gaussian Mixture Model further improves the efficiency and the performance of the method, which we name GMAC. We empirically show that GMAC captures the correct representation of value distributions and improves the performance of a conventional actor-critic method with low computational cost, in both discrete and continuous action spaces using Arcade Learning Environment (ALE) and PyBullet environment.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Trust Region-Based Safe Distributional Reinforcement Learning for Multiple ConstraintsDohyeong Kim, Kyungjae Lee, Songhwai OhNeurIPS 2023 · 被引用 26 次
- The Nature of Temporal Difference Errors in Multi-step Distributional Reinforcement LearningYunhao Tang, Rémi Munos, Mark Rowland, Bernardo Ávila Pires 等NeurIPS 2022 · 被引用 16 次
- Quantile Credit AssignmentThomas Mesnard, Wenqi Chen, Alaa Saade, Yunhao Tang 等ICML 2023 · 被引用 3 次
- Distributional Active InferenceAbdullah Akgül, Gulcin Baykal, Manuel Haussmann, Mustafa Mert Çelikok 等ICML 2026
它引用的顶会 Paper1
相关 Paper
- CTD4 - a Deep Continuous Distributional Actor-Critic Agent with a Kalman Fusion of Multiple CriticsDavid Valencia, Henry Williams, Yuning Xing, Trevor Gee 等AAAI 2025 · 被引用 6 次
- Conjugated Discrete Distributions for Distributional Reinforcement LearningBjörn Lindenberg, Jonas Nordqvist, Karl-Olof LindahlAAAI 2022 · 被引用 2 次
- Distributions as Actions: A Unified Framework for Diverse Action SpacesJiamin He, A. Rupam Mahmood, Martha WhiteICLR 2026
- Sample-Efficient Reinforcement Learning via Conservative Model-Based Actor-CriticZhihai Wang, Jie Wang, Qi Zhou, Bin Li 等AAAI 2022 · 被引用 38 次
- RAMAC: Multimodal Risk-Aware Offline Reinforcement Learning and the Role of Behavior RegularizationKai Fukazawa, Kunal Mundada, Iman SoltaniICML 2026
