RLx2: Training a Sparse Deep Reinforcement Learning Model from Scratch
Yiqin Tan, Pihe Hu, Ling Pan, Jiatai Huang, Longbo Huang
摘要
Training deep reinforcement learning (DRL) models usually requires high computation costs. Therefore, compressing DRL models possesses immense potential for training acceleration and model deployment. However, existing methods that generate small models mainly adopt the knowledge distillation-based approach by iteratively training a dense network. As a result, the training process still demands massive computing resources. Indeed, sparse training from scratch in DRL has not been well explored and is particularly challenging due to non-stationarity in bootstrap training. In this work, we propose a novel sparse DRL training framework, "the Rigged Reinforcement Learning Lottery" (RLx2), which builds upon gradient-based topology evolution and is capable of training a sparse DRL model based entirely on a sparse network. Specifically, RLx2 introduces a novel multi-step TD target mechanism with a dynamic-capacity replay buffer to achieve robust value learning and efficient topology exploration in sparse models. It also reaches state-of-the-art sparse training performance in several tasks, showing 7.5-20model compression with less than 3% performance degradation and up to 20and 50FLOPs reduction for training and inference, respectively.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- The Dormant Neuron Phenomenon in Deep Reinforcement LearningGhada Sokar, Rishabh Agarwal, Pablo Samuel Castro, Utku EvciICML 2023 · 被引用 153 次
- Mixtures of Experts Unlock Parameter Scaling for Deep RLJohan S. Obando-Ceron, Ghada Sokar, Timon Willi, Clare Lyle 等ICML 2024 · 被引用 74 次
- Reinforcement Learning Finetunes Small Subnetworks in Large Language ModelsSagnik Mukherjee, Lifan Yuan, Dilek Hakkani-Tur, Hao PengNeurIPS 2025 · 被引用 43 次
- In value-based deep reinforcement learning, a pruned network is a good networkJohan S. Obando-Ceron, Aaron C. Courville, Pablo Samuel CastroICML 2024 · 被引用 36 次
- Asymmetric Proximal Policy Optimization: mini-critics boost LLM reasoningJiashun Liu, Johan S. Obando-Ceron, Han Lu, Yancheng He 等ICLR 2026 · 被引用 11 次
它引用的顶会 Paper9
- Picking Winning Tickets Before Training by Preserving Gradient FlowChaoqi Wang, Guodong Zhang, Roger B. GrosseICLR 2020 · 被引用 743 次
- Rigging the Lottery: Making All Tickets WinnersUtku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro 等ICML 2020 · 被引用 723 次
- The Lottery Ticket Hypothesis for Pre-trained BERT NetworksTianlong Chen, Jonathan Frankle, Shiyu Chang, Sijia Liu 等NeurIPS 2020 · 被引用 428 次
- Revisiting Fundamentals of Experience ReplayWilliam Fedus, Prajit Ramachandran, Rishabh Agarwal, Yoshua Bengio 等ICML 2020 · 被引用 303 次
- A Unified Lottery Ticket Hypothesis for Graph Neural NetworksTianlong Chen, Yongduo Sui, Xuxi Chen, Aston Zhang 等ICML 2021 · 被引用 208 次
相关 Paper
- Value-Based Deep Multi-Agent Reinforcement Learning with Dynamic Sparse TrainingPihe Hu, Shaolong Li, Zhuoran Li, Ling Pan 等NeurIPS 2024 · 被引用 2 次
- A Unified Self-Regulating Training Framework for Federated Deep Reinforcement LearningMeng Xu, Xinhong Chen, Zhongying Chen, Guanyi Zhao 等AAAI 2026
- Sparse-RL: Breaking the Memory Wall in LLM Reinforcement Learning via Stable Sparse RolloutsSijia Luo, Xiaokang Zhang, Yuxuan Hu, Bohan Zhang 等ACL 2026 · 被引用 6 次
- M3OFF: Module-Compositional Model-Free Computation Offloading in Multi-Environment MECTao Ren, Zheyuan Hu, Jianwei Niu, Weikun Feng 等INFOCOM 2024 · 被引用 5 次
- DECORE: Deep Compression with Reinforcement LearningManoj Alwani, Yang Wang, Vashisht MadhavanCVPR 2022 · 被引用 42 次
