Convergent and Efficient Deep Q Learning Algorithm
Zhikang T. Wang, Masahito Ueda
摘要
Despite the empirical success of the deep Q network (DQN) reinforcement learning algorithm and its variants, DQN is still not well understood and it does not guarantee convergence. In this work, we show that DQN can indeed diverge and cease to operate in realistic settings. Although there exist gradient-based convergent methods, we show that they actually have inherent problems in learning dynamics which cause them to fail even for simple tasks. To overcome these problems, we propose a convergent DQN algorithm (C-DQN) that is guaranteed to converge and can work with large discount factors (0.9998). It learns robustly in difficult settings and can learn several difficult games in the Atari 2600 benchmark that DQN fails to solve.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- Aligning Large Language Models with Representation Editing: A Control PerspectiveLingkai Kong, Haorui Wang, Wenhao Mu, Yuanqi Du 等NeurIPS 2024 · 被引用 80 次
- Temporal Difference Learning: Why It Can Be Fast and How It Will Be FasterPatrick Schnell, Luca Guastoni, Nils ThuereyICLR 2025
相关 Paper
- Self-correcting Q-learningRong Zhu, Mattia RigottiAAAI 2021 · 被引用 22 次
- Weakly Coupled Deep Q-NetworksIbrahim El Shar, Daniel R. JiangNeurIPS 2023 · 被引用 12 次
- Gradient Temporal-Difference Learning with Regularized CorrectionsSina Ghiassian, Andrew Patterson, Shivam Garg, Dhawal Gupta 等ICML 2020 · 被引用 49 次
- Confounding Robust Deep Reinforcement Learning: A Causal ApproachMingxuan Li, Junzhe Zhang, Elias BareinboimNeurIPS 2025 · 被引用 7 次
- Learning Dynamics and Generalization in Deep Reinforcement LearningClare Lyle, Mark Rowland, Will Dabney, Marta Kwiatkowska 等ICML 2022 · 被引用 40 次
