VA-learning as a more efficient alternative to Q-learning
Yunhao Tang, Rémi Munos, Mark Rowland, Michal Valko
摘要
In reinforcement learning, the advantage function is critical for policy improvement, but is often extracted from a learned Q-function. A natural question is: Why not learn the advantage function directly? In this work, we introduce VA-learning, which directly learns advantage function and value function using bootstrapping, without explicit reference to Q-functions. VA-learning learns off-policy and enjoys similar theoretical guarantees as Q-learning. Thanks to the direct learning of advantage function and value function, VA-learning improves the sample efficiency over Q-learning both in tabular implementations and deep RL agents on Atari-57 games. We also identify a close connection between VA-learning and the dueling architecture, which partially explains why a simple architectural change to DQN agents tends to improve performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Action Gaps and Advantages in Continuous-Time Distributional Reinforcement LearningHarley Wiltzer, Marc G. Bellemare, David Meger, Patrick Shafto 等NeurIPS 2024 · 被引用 9 次
- Knowing When to Quit: A Principled Framework for Dynamic Abstention in LLM ReasoningHen Davidov, Nachshon Cohen, Oren Kalinsky, Yaron Fairstein 等ICML 2026 · 被引用 3 次
- Accelerating Q-learning through Efficient Value-sharing across ActionsPrabhat Nagarajan, Brett Daley, Martha White, Marlos C. MachadoICML 2026
它引用的顶会 Paper1
相关 Paper
- Direct Advantage EstimationHsiao-Ru Pan, Nico Gürtler, Alexander Neitz, Bernhard SchölkopfNeurIPS 2022 · 被引用 20 次
- Skill or Luck? Return Decomposition via Advantage FunctionsHsiao-Ru Pan, Bernhard SchölkopfICLR 2024 · 被引用 7 次
- Efficient Off-Policy Learning for High-Dimensional Action SpacesFabian Otto, Philipp Becker, Ngo Anh Vien, Gerhard NeumannICLR 2025
- Munchausen Reinforcement LearningNino Vieillard, Olivier Pietquin, Matthieu GeistNeurIPS 2020 · 被引用 120 次
- Discovering Reinforcement Learning AlgorithmsJunhyuk Oh, Matteo Hessel, Wojciech M. Czarnecki, Zhongwen Xu 等NeurIPS 2020 · 被引用 154 次
