Simplifying Deep Temporal Difference Learning
Matteo Gallici, Mattie Fellows, Benjamin Ellis, Bartomeu Pou, Ivan Masmitja, Jakob Nicolaus Foerster, Mario Martin
摘要
Q-learning played a foundational role in the field reinforcement learning (RL). However, TD algorithms with off-policy data, such as Q-learning, or nonlinear function approximation like deep neural networks require several additional tricks to stabilise training, primarily a large replay buffer and target networks. Unfortunately, the delayed updating of frozen network parameters in the target network harms the sample efficiency and, similarly, the large replay buffer introduces memory and implementation overheads. In this paper, we investigate whether it is possible to accelerate and simplify off-policy TD training while maintaining its stability. Our key theoretical result demonstrates for the first time that regularisation techniques such as LayerNorm can yield provably convergent TD algorithms without the need for a target network or replay buffer, even with off-policy data. Empirically, we find that online, parallelised sampling enabled by vectorised environments stabilises training without the need for a large replay buffer. Motivated by these findings, we propose PQN, our simplified deep online Q-Learning algorithm. Surprisingly, this simple algorithm is competitive with more complex methods like: Rainbow in Atari, PPO-RNN in Craftax, QMix in Smax, and can be up to 50x faster than traditional DQN without sacrificing sample efficiency. In an era where PPO has become the go-to RL algorithm, PQN reestablishes off-policy Q-learning as a viable alternative. We open-source our code at: https://github.com/mttga/purejaxql .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper36
- Horizon Reduction Makes RL ScalableSeohong Park, Kevin Frans, Deepinder Mann, Benjamin Eysenbach 等NeurIPS 2025 · 被引用 60 次
- Stable Gradients for Stable Learning at Scale in Deep Reinforcement LearningRoger Creus Castanyer, Johan S. Obando-Ceron, Lu Li, Pierre-Luc Bacon 等NeurIPS 2025 · 被引用 26 次
- Chunking the Critic: A Transformer-based Soft Actor-Critic with N-Step ReturnsDong Tian, Onur Celik, Gerhard NeumannICLR 2026 · 被引用 18 次
- Can Learned Optimization Make Reinforcement Learning Less Difficult?Alexander David Goldie, Chris Lu, Matthew Thomas Jackson, Shimon Whiteson 等NeurIPS 2024 · 被引用 18 次
- Evolution Strategies at the HyperscaleBidipta Sarkar, Mattie Fellows, Juan Duque, Alistair Letcher 等ICML 2026 · 被引用 16 次
它引用的顶会 Paper19
- Agent57: Outperforming the Atari Human BenchmarkAdrià Puigdomènech Badia, Bilal Piot, Steven Kapturowski, Pablo Sprechmann 等ICML 2020 · 被引用 584 次
- The NetHack Learning EnvironmentHeinrich Küttler, Nantas Nardelli, Alexander H. Miller, Roberta Raileanu 等NeurIPS 2020 · 被引用 251 次
- Behaviour Suite for Reinforcement LearningIan Osband, Yotam Doron, Matteo Hessel, John Aslanides 等ICLR 2020 · 被引用 204 次
- Benchmarking the Spectrum of Agent CapabilitiesDanijar HafnerICLR 2022 · 被引用 193 次
- On the linearity of large non-linear models: when and why the tangent kernel is constantChaoyue Liu, Libin Zhu, Mikhail BelkinNeurIPS 2020 · 被引用 183 次
相关 Paper
- Faster Deep Reinforcement Learning with Slower Online NetworkKavosh Asadi, Rasool Fakoor, Omer Gottesman, Taesup Kim 等NeurIPS 2022 · 被引用 7 次
- Parallel Q-Learning: Scaling Off-policy Reinforcement Learning under Massively Parallel SimulationZechu Li, Tao Chen, Zhang-Wei Hong, Anurag Ajay 等ICML 2023 · 被引用 27 次
- Munchausen Reinforcement LearningNino Vieillard, Olivier Pietquin, Matthieu GeistNeurIPS 2020 · 被引用 120 次
- Bridging the performance-gap between target-free and target-based reinforcement learningThéo Vincent, Yogesh Tripathi, Tim Lukas Faust, Abdullah Akgül 等ICLR 2026 · 被引用 6 次
- Breaking the Deadly Triad with a Target NetworkShangtong Zhang, Hengshuai Yao, Shimon WhitesonICML 2021 · 被引用 61 次
