Direct Advantage Estimation
Hsiao-Ru Pan, Nico Gürtler, Alexander Neitz, Bernhard Schölkopf
摘要
The predominant approach in reinforcement learning is to assign credit to actions based on the expected return. However, we show that the return may depend on the policy in a way which could lead to excessive variance in value estimation and slow down learning. Instead, we show that the advantage function can be interpreted as causal effects and shares similar properties with causal representations. Based on this insight, we propose Direct Advantage Estimation (DAE), a novel method that can model the advantage function and estimate it directly from on-policy data while simultaneously minimizing the variance of the return without requiring the (action-)value function. We also relate our method to Temporal Difference methods by showing how value functions can be seamlessly integrated into DAE. The proposed method is easy to implement and can be readily adapted by modern actor-critic methods. We evaluate DAE empirically on three discrete control domains and show that it can outperform generalized advantage estimation (GAE), a strong baseline for advantage estimation, on a majority of the environments when applied to policy optimization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Action Gaps and Advantages in Continuous-Time Distributional Reinforcement LearningHarley Wiltzer, Marc G. Bellemare, David Meger, Patrick Shafto 等NeurIPS 2024 · 被引用 9 次
- Skill or Luck? Return Decomposition via Advantage FunctionsHsiao-Ru Pan, Bernhard SchölkopfICLR 2024 · 被引用 7 次
- Decision-Aware Actor-Critic with Function Approximation and Theoretical GuaranteesSharan Vaswani, Amirreza Kazemi, Reza Babanezhad Harikandeh, Nicolas Le RouxNeurIPS 2023 · 被引用 6 次
- Relative Value LearningMarc Höftmann, Jan Robine, Stefan HarmelingICLR 2026 · 被引用 1 次
- On the Adversarial Robustness of Multi-Kernel ClusteringHao Yu, Weixuan Liang, Ke Liang, Suyuan Liu 等ICML 2025
它引用的顶会 Paper5
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Batch Normalization Biases Residual Blocks Towards the Identity Function in Deep NetworksSoham De, Samuel L. SmithNeurIPS 2020 · 被引用 173 次
- Off-Policy Evaluation in Partially Observable EnvironmentsGuy Tennenholtz, Uri Shalit, Shie MannorAAAI 2020 · 被引用 91 次
- Counterfactual Credit Assignment in Model-Free Reinforcement LearningThomas Mesnard, Theophane Weber, Fabio Viola, Shantanu Thakoor 等ICML 2021 · 被引用 70 次
- Sequential Causal Imitation Learning with Unobserved ConfoundersDaniel Kumor, Junzhe Zhang, Elias BareinboimNeurIPS 2021 · 被引用 53 次
相关 Paper
- VA-learning as a more efficient alternative to Q-learningYunhao Tang, Rémi Munos, Mark Rowland, Michal ValkoICML 2023 · 被引用 11 次
- Robust Action Gap Increasing with Clipped Advantage LearningZhe Zhang, Yaozhong Gan, Xiaoyang TanAAAI 2022 · 被引用 3 次
- Decoupling Value and Policy for Generalization in Reinforcement LearningRoberta Raileanu, Rob FergusICML 2021 · 被引用 116 次
- Value-driven Hindsight ModellingArthur Guez, Fabio Viola, Theophane Weber, Lars Buesing 等NeurIPS 2020 · 被引用 12 次
- Smoothing Advantage LearningYaozhong Gan, Zhe Zhang, Xiaoyang TanAAAI 2022 · 被引用 3 次
