Value-Based Deep Multi-Agent Reinforcement Learning with Dynamic Sparse Training
Pihe Hu, Shaolong Li, Zhuoran Li, Ling Pan, Longbo Huang
摘要
Deep Multi-agent Reinforcement Learning (MARL) relies on neural networks with numerous parameters in multi-agent scenarios, often incurring substantial computational overhead. Consequently, there is an urgent need to expedite training and enable model compression in MARL. This paper proposes the utilization of dynamic sparse training (DST), a technique proven effective in deep supervised learning tasks, to alleviate the computational burdens in MARL training. However, a direct adoption of DST fails to yield satisfactory MARL agents, leading to breakdowns in value learning within deep sparse value-based MARL models. Motivated by this challenge, we introduce an innovative Multi-Agent Sparse Training (MAST) framework aimed at simultaneously enhancing the reliability of learning targets and the rationality of sample distribution to improve value learning in sparse models. Specifically, MAST incorporates the Soft Mellowmax Operator with a hybrid TD-() schema to establish dependable learning targets. Additionally, it employs a dual replay buffer mechanism to enhance the distribution of training samples. Building upon these aspects, MAST utilizes gradient-based topology evolution to exclusively train multiple MARL agents using sparse networks. Our comprehensive experimental investigation across various value-based MARL algorithms on multiple benchmarks demonstrates, for the first time, significant reductions in redundancy of up to in Floating Point Operations (FLOPs) for both training and inference, with less than performance degradation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper18
- Weighted QMIX: Expanding Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement LearningTabish Rashid, Gregory Farquhar, Bei Peng, Shimon WhitesonNeurIPS 2020 · 被引用 1,960 次
- Emergent Tool Use From Multi-Agent AutocurriculaBowen Baker, Ingmar Kanitscheider, Todor M. Markov, Yi Wu 等ICLR 2020 · 被引用 751 次
- Picking Winning Tickets Before Training by Preserving Gradient FlowChaoqi Wang, Guodong Zhang, Roger B. GrosseICLR 2020 · 被引用 743 次
- Rigging the Lottery: Making All Tickets WinnersUtku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro 等ICML 2020 · 被引用 723 次
- The Lottery Ticket Hypothesis for Pre-trained BERT NetworksTianlong Chen, Jonathan Frankle, Shiyu Chang, Sijia Liu 等NeurIPS 2020 · 被引用 428 次
相关 Paper
- RLx2: Training a Sparse Deep Reinforcement Learning Model from ScratchYiqin Tan, Pihe Hu, Ling Pan, Jiatai Huang 等ICLR 2023 · 被引用 1 次
- A Unified Self-Regulating Training Framework for Federated Deep Reinforcement LearningMeng Xu, Xinhong Chen, Zhongying Chen, Guanyi Zhao 等AAAI 2026
- MASER: Multi-Agent Reinforcement Learning with Subgoals Generated from Experience Replay BufferJeewon Jeon, Woojun Kim, Whiyoung Jung, Youngchul SungICML 2022 · 被引用 53 次
- Doubly Robust Augmented Transfer for Meta-Reinforcement LearningYuankun Jiang, Nuowen Kan, Chenglin Li, Wenrui Dai 等NeurIPS 2023 · 被引用 3 次
- The State of Sparse Training in Deep Reinforcement LearningLaura Graesser, Utku Evci, Erich Elsen, Pablo Samuel CastroICML 2022 · 被引用 65 次
