Simplifying Deep Temporal Difference Learning
Matteo Gallici, Mattie Fellows, Benjamin Ellis, Bartomeu Pou, Ivan Masmitja, Jakob Nicolaus Foerster, Mario Martin
Abstract
Q-learning played a foundational role in the field reinforcement learning (RL). However, TD algorithms with off-policy data, such as Q-learning, or nonlinear function approximation like deep neural networks require several additional tricks to stabilise training, primarily a large replay buffer and target networks. Unfortunately, the delayed updating of frozen network parameters in the target network harms the sample efficiency and, similarly, the large replay buffer introduces memory and implementation overheads. In this paper, we investigate whether it is possible to accelerate and simplify off-policy TD training while maintaining its stability. Our key theoretical result demonstrates for the first time that regularisation techniques such as LayerNorm can yield provably convergent TD algorithms without the need for a target network or replay buffer, even with off-policy data. Empirically, we find that online, parallelised sampling enabled by vectorised environments stabilises training without the need for a large replay buffer. Motivated by these findings, we propose PQN, our simplified deep online Q-Learning algorithm. Surprisingly, this simple algorithm is competitive with more complex methods like: Rainbow in Atari, PPO-RNN in Craftax, QMix in Smax, and can be up to 50x faster than traditional DQN without sacrificing sample efficiency. In an era where PPO has become the go-to RL algorithm, PQN reestablishes off-policy Q-learning as a viable alternative. We open-source our code at: https://github.com/mttga/purejaxql .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0e104f69-ad45-4e7b-b4d9-262bc0f17f8bCited by top-tier papers36
- Horizon Reduction Makes RL ScalableSeohong Park, Kevin Frans, Deepinder Mann, Benjamin Eysenbach et al.NeurIPS 2025 · 60 citations
- Stable Gradients for Stable Learning at Scale in Deep Reinforcement LearningRoger Creus Castanyer, Johan S. Obando-Ceron, Lu Li, Pierre-Luc Bacon et al.NeurIPS 2025 · 26 citations
- Chunking the Critic: A Transformer-based Soft Actor-Critic with N-Step ReturnsDong Tian, Onur Celik, Gerhard NeumannICLR 2026 · 18 citations
- Can Learned Optimization Make Reinforcement Learning Less Difficult?Alexander David Goldie, Chris Lu, Matthew Thomas Jackson, Shimon Whiteson et al.NeurIPS 2024 · 18 citations
- Evolution Strategies at the HyperscaleBidipta Sarkar, Mattie Fellows, Juan Duque, Alistair Letcher et al.ICML 2026 · 16 citations
Builds on19
- Agent57: Outperforming the Atari Human BenchmarkAdrià Puigdomènech Badia, Bilal Piot, Steven Kapturowski, Pablo Sprechmann et al.ICML 2020 · 584 citations
- The NetHack Learning EnvironmentHeinrich Küttler, Nantas Nardelli, Alexander H. Miller, Roberta Raileanu et al.NeurIPS 2020 · 251 citations
- Behaviour Suite for Reinforcement LearningIan Osband, Yotam Doron, Matteo Hessel, John Aslanides et al.ICLR 2020 · 204 citations
- Benchmarking the Spectrum of Agent CapabilitiesDanijar HafnerICLR 2022 · 193 citations
- On the linearity of large non-linear models: when and why the tangent kernel is constantChaoyue Liu, Libin Zhu, Mikhail BelkinNeurIPS 2020 · 183 citations
Related papers
- Faster Deep Reinforcement Learning with Slower Online NetworkKavosh Asadi, Rasool Fakoor, Omer Gottesman, Taesup Kim et al.NeurIPS 2022 · 7 citations
- Parallel Q-Learning: Scaling Off-policy Reinforcement Learning under Massively Parallel SimulationZechu Li, Tao Chen, Zhang-Wei Hong, Anurag Ajay et al.ICML 2023 · 27 citations
- Munchausen Reinforcement LearningNino Vieillard, Olivier Pietquin, Matthieu GeistNeurIPS 2020 · 120 citations
- Bridging the performance-gap between target-free and target-based reinforcement learningThéo Vincent, Yogesh Tripathi, Tim Lukas Faust, Abdullah Akgül et al.ICLR 2026 · 6 citations
- Breaking the Deadly Triad with a Target NetworkShangtong Zhang, Hengshuai Yao, Shimon WhitesonICML 2021 · 61 citations
