Direct Advantage Estimation
Hsiao-Ru Pan, Nico Gürtler, Alexander Neitz, Bernhard Schölkopf
Abstract
The predominant approach in reinforcement learning is to assign credit to actions based on the expected return. However, we show that the return may depend on the policy in a way which could lead to excessive variance in value estimation and slow down learning. Instead, we show that the advantage function can be interpreted as causal effects and shares similar properties with causal representations. Based on this insight, we propose Direct Advantage Estimation (DAE), a novel method that can model the advantage function and estimate it directly from on-policy data while simultaneously minimizing the variance of the return without requiring the (action-)value function. We also relate our method to Temporal Difference methods by showing how value functions can be seamlessly integrated into DAE. The proposed method is easy to implement and can be readily adapted by modern actor-critic methods. We evaluate DAE empirically on three discrete control domains and show that it can outperform generalized advantage estimation (GAE), a strong baseline for advantage estimation, on a majority of the environments when applied to policy optimization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e5bec354-c5d9-42d2-8e48-a6357f1e627bCited by top-tier papers5
- Action Gaps and Advantages in Continuous-Time Distributional Reinforcement LearningHarley Wiltzer, Marc G. Bellemare, David Meger, Patrick Shafto et al.NeurIPS 2024 · 9 citations
- Skill or Luck? Return Decomposition via Advantage FunctionsHsiao-Ru Pan, Bernhard SchölkopfICLR 2024 · 7 citations
- Decision-Aware Actor-Critic with Function Approximation and Theoretical GuaranteesSharan Vaswani, Amirreza Kazemi, Reza Babanezhad Harikandeh, Nicolas Le RouxNeurIPS 2023 · 6 citations
- Relative Value LearningMarc Höftmann, Jan Robine, Stefan HarmelingICLR 2026 · 1 citation
- On the Adversarial Robustness of Multi-Kernel ClusteringHao Yu, Weixuan Liang, Ke Liang, Suyuan Liu et al.ICML 2025
Builds on5
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Batch Normalization Biases Residual Blocks Towards the Identity Function in Deep NetworksSoham De, Samuel L. SmithNeurIPS 2020 · 173 citations
- Off-Policy Evaluation in Partially Observable EnvironmentsGuy Tennenholtz, Uri Shalit, Shie MannorAAAI 2020 · 91 citations
- Counterfactual Credit Assignment in Model-Free Reinforcement LearningThomas Mesnard, Theophane Weber, Fabio Viola, Shantanu Thakoor et al.ICML 2021 · 70 citations
- Sequential Causal Imitation Learning with Unobserved ConfoundersDaniel Kumor, Junzhe Zhang, Elias BareinboimNeurIPS 2021 · 53 citations
Related papers
- VA-learning as a more efficient alternative to Q-learningYunhao Tang, Rémi Munos, Mark Rowland, Michal ValkoICML 2023 · 11 citations
- Robust Action Gap Increasing with Clipped Advantage LearningZhe Zhang, Yaozhong Gan, Xiaoyang TanAAAI 2022 · 3 citations
- Decoupling Value and Policy for Generalization in Reinforcement LearningRoberta Raileanu, Rob FergusICML 2021 · 116 citations
- Value-driven Hindsight ModellingArthur Guez, Fabio Viola, Theophane Weber, Lars Buesing et al.NeurIPS 2020 · 12 citations
- Smoothing Advantage LearningYaozhong Gan, Zhe Zhang, Xiaoyang TanAAAI 2022 · 3 citations
