SrSv: Integrating Sequential Rollouts with Sequential Value Estimation for Multi-agent Reinforcement Learning
Xu Wan, Chao Yang, Cheng Yang, Jie Song, Mingyang Sun
Abstract
Although multi-agent reinforcement learning (MARL) has shown its success across diverse domains, extending its application to large-scale real-world systems still faces significant challenges. Primarily, the high complexity of real-world environments exacerbates the credit assignment problem, substantially reducing training efficiency. Moreover, the variability of agent populations in large-scale scenarios necessitates scalable decision-making mechanisms. To address these challenges, we propose a novel framework: Sequential rollout with Sequential value estimation (SrSv). This framework aims to capture agent interdependence and provide a scalable solution for cooperative MARL. Specifically, SrSv leverages the autoregressive property of the Transformer model to handle varying populations through sequential action rollout. Furthermore, to capture the interdependence of policy distributions and value functions among multiple agents, we introduce an innovative sequential value estimation methodology and integrates the value approximation into an attention-based sequential model.
We evaluate SrSv on three benchmarks: Multi-Agent MuJoCo, StarCraft Multi-Agent Challenge, and DubinsCars. Experimental results demonstrate that SrSv significantly outperforms baseline methods in terms of training efficiency without compromising convergence performance. Moreover, when implemented in a large-scale DubinsCar system with 1,024 agents, our framework surpasses existing benchmarks, highlighting the excellent scalability of SrSv.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0bfeee59-ddec-4d41-8e8f-c35dd11a1035Cited by top-tier papers2
- HCPO: Hierarchical Conductor-Based Policy Optimization in Multi-Agent Reinforcement LearningZejiao Liu, Junqi Tu, Yitian Hong, Luolin Xiong et al.AAAI 2026
- M2I2: Learning Efficient Multi-Agent Communication via Masked State Modeling and Intention InferenceChuxiong Sun, Peng He, Qirui Ji, Zehua Zang et al.AAAI 2026
Builds on10
- QPLEX: Duplex Dueling Multi-Agent Q-LearningJianhao Wang, Zhizhou Ren, Terry Liu, Yang Yu et al.ICLR 2021 · 595 citations
- Trust Region Policy Optimisation in Multi-Agent Reinforcement LearningJakub Grudzien Kuba, Ruiqing Chen, Muning Wen, Ying Wen et al.ICLR 2022 · 367 citations
- Settling the Variance of Multi-Agent Policy GradientsJakub Grudzien Kuba, Muning Wen, Linghui Meng, Shangding Gu et al.NeurIPS 2021 · 121 citations
- FOP: Factorizing Optimal Joint Policy of Maximum-Entropy Multi-Agent Reinforcement LearningTianhao Zhang, Yueheng Li, Chen Wang, Guangming Xie et al.ICML 2021 · 88 citations
- Q-value Path Decomposition for Deep Multiagent Reinforcement LearningYaodong Yang, Jianye Hao, Guangyong Chen, Hongyao Tang et al.ICML 2020 · 64 citations
Related papers
- Sequential Asynchronous Action Coordination in Multi-Agent Systems: A Stackelberg Decision Transformer ApproachBin Zhang, Hangyu Mao, Lijuan Li, Zhiwei Xu et al.ICML 2024 · 13 citations
- Multi-Agent Reinforcement Learning is a Sequence Modeling ProblemMuning Wen, Jakub Grudzien Kuba, Runji Lin, Weinan Zhang et al.NeurIPS 2022 · 408 citations
- Sable: a Performant, Efficient and Scalable Sequence Model for MARLOmayma Mahjoub, Sasha Abramowitz, Ruan John de Kock, Wiem Khlifi et al.ICML 2025
- Shapley Counterfactual Credits for Multi-Agent Reinforcement LearningJiahui Li, Kun Kuang, Baoxiang Wang, Furui Liu et al.KDD 2021 · 49 citations
- High-order Interactions Modeling for Interpretable Multi-Agent Q-LearningQinyu Xu, Yuanyang Zhu, Xuefei Wu, Chunlin ChenNeurIPS 2025 · 2 citations
