Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers
Gautham Vasan, Mohamed Elsayed, Seyed Alireza Azimi, Jiamin He, Fahim Shahriar, Colin Bellinger, Martha White, Rupam Mahmood
Abstract
Modern deep policy gradient methods achieve effective performance on simulated robotic tasks, but they all require large replay buffers or expensive batch updates, or both, making them incompatible for real systems with resource-limited computers. We show that these methods fail catastrophically when limited to small replay buffers or during incremental learning, where updates only use the most recent sample without batch updates or a replay buffer. We propose a novel incremental deep policy gradient method -- Action Value Gradient (AVG) and a set of normalization and scaling techniques to address the challenges of instability in incremental learning. On robotic simulation benchmarks, we show that AVG is the only incremental method that learns effectively, often achieving final performance comparable to batch policy gradient methods. This advancement enabled us to show for the first time effective deep reinforcement learning with real robots using only incremental updates, employing a robotic manipulator and a mobile robot.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Bridging the performance-gap between target-free and target-based reinforcement learningThéo Vincent, Yogesh Tripathi, Tim Lukas Faust, Abdullah Akgül et al.ICLR 2026 · 6 citations
- Distributions as Actions: A Unified Framework for Diverse Action SpacesJiamin He, A. Rupam Mahmood, Martha WhiteICLR 2026
- Squeezing More from the Stream : Learning Representation Online for Streaming Reinforcement LearningNilaksh, Antoine Clavaud, Mathieu Reymond, Francois Rivest et al.ICML 2026
- Network Sparsity Unlocks the Scaling Potential of Deep Reinforcement LearningGuozheng Ma, Lu Li, Zilin Wang, Li Shen et al.ICML 2025
Builds on14
- On the Global Convergence Rates of Softmax Policy Gradient MethodsJincheng Mei, Chenjun Xiao, Csaba Szepesvári, Dale SchuurmansICML 2020 · 349 citations
- Implementation Matters in Deep RL: A Case Study on PPO and TRPOLogan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras et al.ICLR 2020 · 305 citations
- The Primacy Bias in Deep Reinforcement LearningEvgenii Nikishin, Max Schwarzer, Pierluca D'Oro, Pierre-Luc Bacon et al.ICML 2022 · 269 citations
- Understanding Plasticity in Neural NetworksClare Lyle, Zeyu Zheng, Evgenii Nikishin, Bernardo Ávila Pires et al.ICML 2023 · 162 citations
- Bigger, Better, Faster: Human-level Atari with human-level efficiencyMax Schwarzer, Johan S. Obando-Ceron, Aaron C. Courville, Marc G. Bellemare et al.ICML 2023 · 155 citations
Related papers
- Sim and Real: Better TogetherShirli Di-Castro Shashua, Dotan Di Castro, Shie MannorNeurIPS 2021 · 14 citations
- iManip: Skill-Incremental Learning for Robotic ManipulationZexin Zheng, Jia-Feng Cai, Xiao-Ming Wu, Yi-Lin Wei et al.ICCV 2025 · 14 citations
- Smart Replay: Adaptive Scheduling of Memory Rehearsal for Computational Resource-Aware Incremental LearningJianting Chen, Dianzhi Yu, Irwin KingCVPR 2026
- Uncertainty-Aware Policy Optimization: A Robust, Adaptive Trust Region ApproachJames Queeney, Ioannis Ch. Paschalidis, Christos G. CassandrasAAAI 2021 · 11 citations
- Incremental Reinforcement Learning with Dual-Adaptive ε-Greedy ExplorationWei Ding, Siyang Jiang, Hsi-Wen Chen, Ming-Syan ChenAAAI 2023 · 11 citations
