MAD-TD: Model-Augmented Data stabilizes High Update Ratio RL
Claas Voelcker, Marcel Hussing, Eric Eaton, Amir-massoud Farahmand, Igor Gilitschenski
Abstract
Building deep reinforcement learning (RL) agents that find a good policy with few samples has proven notoriously challenging. To achieve sample efficiency, recent work has explored updating neural networks with large numbers of gradient steps for every new sample. While such high update-to-data (UTD) ratios have shown strong empirical performance, they also introduce instability to the training process. Previous approaches need to rely on periodic neural network parameter resets to address this instability, but restarting the training process is infeasible in many real-world applications and requires tuning the resetting interval. In this paper, we focus on one of the core difficulties of stable training with limited samples: the inability of learned value functions to generalize to unobserved on-policy actions. We mitigate this issue directly by augmenting the off-policy RL training process with a small amount of data generated from a learned world model. Our method, Model-Augmented Data for Temporal Difference learning (MAD-TD) uses small amounts of generated data to stabilize high UTD training and achieve competitive performance on the most challenging tasks in the DeepMind control suite. Our experiments further highlight the importance of employing a good model to generate data, MAD-TD's ability to combat value overestimation, and its practical stability gains for continued learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 84a77af7-93b4-4ae1-ae05-2ce75ac1ae40Cited by top-tier papers6
- Simplicial Embeddings Improve Sample Efficiency in Actor–Critic AgentsJohan Obando-Ceron, Walter Mayor, Samuel Lavoie, Scott Fujimoto et al.ICLR 2026 · 12 citations
- Relative Entropy Pathwise Policy OptimizationClaas Voelcker, Axel Brunnbauer, Marcel Hussing, Michal Nauman et al.ICLR 2026 · 6 citations
- Replicable Reinforcement Learning with Linear Function ApproximationEric Eaton, Marcel Hussing, Michael Kearns, Aaron Roth et al.ICLR 2026 · 6 citations
- Replicable Reinforcement LearningEric Eaton, Marcel Hussing, Michael Kearns, Jessica SorrellNeurIPS 2023 · 3 citations
- R2R2: Robust Representation for Intensive Experience Reuse via Redundancy Reduction in Self-Predictive LearningSanghyeob Song, Donghyeok Lee, Jinsik Kim, Sungroh YoonICML 2026
Builds on47
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 1,170 citations
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville et al.NeurIPS 2021 · 1,067 citations
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon et al.NeurIPS 2020 · 989 citations
Related papers
- Sample-Efficient Reinforcement Learning by Breaking the Replay Ratio BarrierPierluca D'Oro, Max Schwarzer, Evgenii Nikishin, Pierre-Luc Bacon et al.ICLR 2023
- Learning Dynamics and Generalization in Deep Reinforcement LearningClare Lyle, Mark Rowland, Will Dabney, Marta Kwiatkowska et al.ICML 2022 · 40 citations
- Scaling Off-Policy Reinforcement Learning with Batch and Weight NormalizationDaniel Palenicek, Florian Vogt, Joe Watson, Jan PetersNeurIPS 2025 · 22 citations
- Simplified Temporal Consistency Reinforcement LearningYi Zhao, Wenshuai Zhao, Rinu Boney, Juho Kannala et al.ICML 2023 · 19 citations
- Dynamic Update-to-Data Ratio: Minimizing World Model OverfittingNicolai Dorka, Tim Welschehold, Wolfram BurgardICLR 2023 · 1 citation
