Sable: a Performant, Efficient and Scalable Sequence Model for MARL
Omayma Mahjoub, Sasha Abramowitz, Ruan John de Kock, Wiem Khlifi, Simon du Toit, Jemma Daniel, Louay Ben Nessir, Louise Beyers, Juan Claude Formanek, Liam Clark, Arnu Pretorius
Abstract
As multi-agent reinforcement learning (MARL) progresses towards solving larger and more complex problems, it becomes increasingly important that algorithms exhibit the key properties of (1) strong performance, (2) memory efficiency, and (3) scalability. In this work, we introduce Sable, a performant, memory-efficient, and scalable sequence modelling approach to MARL. Sable works by adapting the retention mechanism in Retentive Networks (Sun et al., 2023) to achieve computationally efficient processing of multi-agent observations with long context memory for temporal reasoning. Through extensive evaluations across six diverse environments, we demonstrate how Sable is able to significantly outperform existing state-of-the-art methods in a large number of diverse tasks (34 out of 45 tested). Furthermore, Sable maintains performance as we scale the number of agents, handling environments with more than a thousand agents while exhibiting a linear increase in memory usage. Finally, we conduct ablation studies to isolate the source of Sable's performance gains and confirm its efficient computational memory usage. All experimental data, hyperparameters, and code for a frozen version of Sable used in this paper are available on our website. An improved and maintained version of Sable is available in Mava. * Equal contribution 1 InstaDeep. Correspondence to: Ruan de Kock, Arnu Pretorius <r.dekock, a.pretorius@instadeep.com>.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d79ace7c-721d-4c0c-a2d2-fa27f01c7e1eCited by top-tier papers3
- Multi-Agent Guided Policy OptimizationYueheng Li, Guangming Xie, Zongqing LuICLR 2026 · 4 citations
- Breaking the Performance Ceiling in Reinforcement Learning requires Inference StrategiesFélix Chalumeau, Daniel Rajaonarivonivelomanantsoa, Ruan John de Kock, Juan Claude Formanek et al.NeurIPS 2025 · 2 citations
- Oryx: a Scalable Sequence Model for Many-Agent Coordination in Offline MARLJuan Claude Formanek, Omayma Mahjoub, Louay Ben Nessir, Sasha Abramowitz et al.NeurIPS 2025
Builds on18
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville et al.NeurIPS 2021 · 1,067 citations
- Stabilizing Transformers for Reinforcement LearningEmilio Parisotto, H. Francis Song, Jack W. Rae, Razvan Pascanu et al.ICML 2020 · 464 citations
- Multi-Agent Reinforcement Learning is a Sequence Modeling ProblemMuning Wen, Jakub Grudzien Kuba, Runji Lin, Weinan Zhang et al.NeurIPS 2022 · 408 citations
Related papers
- Multi-agent In-context Coordination via Decentralized Memory RetrievalTao Jiang, Zichuan Lin, Lihe Li, Yi-Chen Li et al.AAAI 2026 · 1 citation
- SrSv: Integrating Sequential Rollouts with Sequential Value Estimation for Multi-agent Reinforcement LearningXu Wan, Chao Yang, Cheng Yang, Jie Song et al.AAAI 2025 · 2 citations
- MARTI: A Framework for Multi-Agent LLM Systems Reinforced Training and InferenceKaiyan Zhang, Kai Tian, Runze Liu, Sihang Zeng et al.ICLR 2026
- RPM: Generalizable Multi-Agent Policies for Multi-Agent Reinforcement LearningWei Qiu, Xiao Ma, Bo An, Svetlana Obraztsova et al.ICLR 2023 · 1 citation
- Learning Multi-Agent Communication through Structured Attentive ReasoningMurtaza Rangwala, Ryan WilliamsNeurIPS 2020 · 42 citations
