DyPS: Dynamic Parameter Sharing in Multi-Agent Reinforcement Learning for Spatio-Temporal Resource Allocation
Jingwei Wang, Qianyue Hao, Wenzhen Huang, Xiaochen Fan, Zhentao Tang, Bin Wang, Jianye Hao, Yong Li
Abstract
In large-scale metropolis, it is critical to efficiently allocate various resources such as electricity, medical care, and transportation to meet the living demands of citizens, according to the spatio-temporal distributions of resources and demands. Previous researchers have done plentiful work on such problems by leveraging Multi-Agent Reinforcement Learning (MARL) methods, where multiple agents cooperatively regulate and allocate the resources to meet the demands. However, facing the great number of agents in large cities, existing MARL methods lack efficient parameter sharing strategies among agents to reduce computational complexity. There remain two primary challenges in efficient parameter sharing: (1) during the RL training process, the behavior of agents changes significantly, limiting the performance of group parameter sharing based on fixed role division decided before training; (2) the behavior of agents forms complicated action trajectories, where their role characteristics are implicit, adding difficulty to dynamically adjusting agent role divisions during the training process. In this paper, we propose Dynamic Parameter Sharing (DyPS) to solve the above challenges. We design self-supervised learning tasks to extract the implicit behavioral characteristics from the action trajectories of agents. Based on the obtained behavioral characteristics, we propose a hierarchical MARL framework capable of dynamically revising the agent role divisions during the training process and thus shares parameters among agents with the same role, reducing computational complexity. In addition, our framework can be combined with various typical MARL algorithms, including IPPO, MAPPO, etc. We conduct 7 experiments in 4 representative resource allocation scenarios, where extensive results demonstrate our method's superior performance, outperforming the state-of-the-art baseline methods by up to 31%. Our source codes are available at https://github.com/tsinghua-fib-lab/DyPS.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers2
- LLM-Explorer: A Plug-in Reinforcement Learning Policy Exploration Enhancement Driven by Large Language ModelsQianyue Hao, Yiwen Song, Qingmin Liao, Jian Yuan et al.NeurIPS 2025 · 6 citations
- Triple-BERT: Do We Really Need MARL for Order Dispatch on Ride-Sharing Platforms?Zijian Zhao, Sen LiICLR 2026 · 4 citations
Related papers
- Towards Complete Multi-Agent Coordination Policy Learning via Denoising Maximum Entropy OptimizationGuanghao Li, lei yuan, Ruiqi Xue, Hengchang Zhang et al.ICML 2026
- Scaling Multi-Agent Reinforcement Learning with Selective Parameter SharingFilippos Christianos, Georgios Papoudakis, Arrasy Rahman, Stefano V. AlbrechtICML 2021 · 165 citations
- LDSA: Learning Dynamic Subtask Assignment in Cooperative Multi-Agent Reinforcement LearningMingyu Yang, Jian Zhao, Xunhan Hu, Wengang Zhou et al.NeurIPS 2022 · 61 citations
- CoopRide: Cooperate All Grids in City-Scale Ride-Hailing Dispatching with Multi-Agent Reinforcement LearningJingwei Wang, Qianyue Hao, Wenzhen Huang, Xiaochen Fan et al.KDD 2025 · 4 citations
- Learning to Share in Networked Multi-Agent Reinforcement LearningYuxuan Yi, Ge Li, Yaowei Wang, Zongqing LuNeurIPS 2022 · 13 citations
