Learning to Simulate Self-driven Particles System with Coordinated Policy Optimization
Zhenghao Peng, Quanyi Li, Ka-Ming Hui, Chunxiao Liu, Bolei Zhou
Abstract
Self-Driven Particles (SDP) describe a category of multi-agent systems common in everyday life, such as flocking birds and traffic flows. In a SDP system, each agent pursues its own goal and constantly changes its cooperative or competitive behaviors with its nearby agents. Manually designing the controllers for such SDP system is time-consuming, while the resulting emergent behaviors are often not realistic nor generalizable. Thus the realistic simulation of SDP systems remains challenging. Reinforcement learning provides an appealing alternative for automating the development of the controller for SDP. However, previous multiagent reinforcement learning (MARL) methods define the agents to be teammates or enemies before hand, which fail to capture the essence of SDP where the role of each agent varies to be cooperative or competitive even within one episode. To simulate SDP with MARL, a key challenge is to coordinate agents' behaviors while still maximizing individual objectives. Taking traffic simulation as the testing bed, in this work we develop a novel MARL method called Coordinated Policy Optimization (CoPO), which incorporates social psychology principle to learn neural controller for SDP. Experiments show that the proposed method can achieve superior performance compared to MARL baselines in various metrics. Noticeably the trained vehicles exhibit complex and diverse social behaviors that improve performance and safety of the population as a whole. Demo video and source code are available at: https://decisionforce.github.io/CoPO/ . When the interactive environment is available, reinforcement learning becomes a promising approach to learn the controllers for actuating the SDP. Recently, many multi-agent reinforcement learning (MARL) methods have been developed to play competitive multi-player games, such as Hide and Seek [1], Football [26] , Go and other board games [41], and StarCraft [40]. However, it is challenging to apply the existing MARL to simulate SDP systems. One essential issue is that 35th Conference on Neural Information Processing Systems (NeurIPS 2021).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- Robust Multi-Agent Coordination via Evolutionary Generation of Auxiliary Adversarial AttackersLei Yuan, Ziqian Zhang, Ke Xue, Hao Yin et al.AAAI 2023 · 31 citations
- Exploring both Individuality and Cooperation for Air-Ground Spatial Crowdsourcing by Multi-Agent Deep Reinforcement LearningYuxiao Ye, Chi Harold Liu, Zipeng Dai, Jianxin Zhao et al.ICDE 2023 · 26 citations
- CoopRide: Cooperate All Grids in City-Scale Ride-Hailing Dispatching with Multi-Agent Reinforcement LearningJingwei Wang, Qianyue Hao, Wenzhen Huang, Xiaochen Fan et al.KDD 2025 · 4 citations
- RPM: Generalizable Multi-Agent Policies for Multi-Agent Reinforcement LearningWei Qiu, Xiao Ma, Bo An, Svetlana Obraztsova et al.ICLR 2023 · 1 citation
- Unreal-MAP: Unreal-Engine-Based General Platform for Multi-agent Reinforcement LearningTianyi Hu, Qingxu Fu, Zhiqiang Pu, Yuan Wang et al.AAAI 2026 · 1 citation
Builds on4
- Emergent Tool Use From Multi-Agent AutocurriculaBowen Baker, Ingmar Kanitscheider, Todor M. Markov, Yi Wu et al.ICLR 2020 · 751 citations
- Google Research Football: A Novel Reinforcement Learning EnvironmentKarol Kurach, Anton Raichuk, Piotr Stanczyk, Michal Zajac et al.AAAI 2020 · 496 citations
- Learning Implicit Credit Assignment for Cooperative Multi-Agent Reinforcement LearningMeng Zhou, Ziyu Liu, Pengwei Sui, Yixuan Li et al.NeurIPS 2020 · 142 citations
- Emergent Road Rules In Multi-Agent Driving EnvironmentsAvik Pal, Jonah Philion, Yuan-Hong Liao, Sanja FidlerICLR 2021 · 15 citations
Related papers
- Scalable Constrained Policy Optimization for Safe Multi-agent Reinforcement LearningLijun Zhang, Lin Li, Wei Wei, Huizhong Song et al.NeurIPS 2024 · 22 citations
- R3DM: Enabling Role Discovery and Diversity Through Dynamics Models in Multi-agent Reinforcement LearningHarsh Goel, Mohammad Omama, Behdad Chalaki, Vaishnav Tadiparthi et al.ICML 2025
- SPACeR: Self-Play Anchoring with Centralized Reference ModelsWei-Jer Chang, Akshay Rangesh, Kevin Joseph, Matthew Strong et al.ICLR 2026 · 9 citations
- Towards Generalizable Multi-Policy Optimization with Self-Evolution for Job SchedulingInguk Choi, Woo-Jin Shin, Sang-Hyun Cho, Hyun-Jung KimNeurIPS 2025 · 4 citations
- Self-Organized Polynomial-Time Coordination GraphsQianlan Yang, Weijun Dong, Zhizhou Ren, Jianhao Wang et al.ICML 2022 · 20 citations
