MAPDP: Cooperative Multi-Agent Reinforcement Learning to Solve Pickup and Delivery Problems
Zefang Zong, Meng Zheng, Yong Li, Depeng Jin
摘要
Cooperative Pickup and Delivery Problem (PDP), as a variant of the typical Vehicle Routing Problems (VRP), is an important formulation in many real-world applications, such as on-demand delivery, industrial warehousing, etc. It is of great importance to efficiently provide high-quality solutions of cooperative PDP. However, it is not trivial to provide effective solutions directly due to two major challenges: 1) the structural dependency between pickup and delivery pairs require explicit modeling and representation. 2) the cooperation between different vehicles is highly related to solution exploration and is difficult to model. In this paper, we propose a novel multi-agent reinforcement learning-based framework to solve the cooperative PDP (MAPDP). First, we design a paired context embedding to well measure the dependency of different nodes considering their structural limits. Second, we utilize cooperative multi-agent decoders to leverage the decision dependence among different vehicle agents based on a special communication embedding. Third, we design a novel cooperative A2C algorithm to train the integrated model. We conduct extensive experiments on a randomly generated dataset and a real-world dataset. Experiments result shown that the proposed MAPDP outperforms all other baselines by at least 1.64% in all settings, and shows significant computation speed during solution inference.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- DeepACO: Neural-enhanced Ant Systems for Combinatorial OptimizationHaoran Ye, Jiarui Wang, Zhiguang Cao, Helan Liang 等NeurIPS 2023 · 被引用 158 次
- Distilling Autoregressive Models to Obtain High-Performance Non-autoregressive Solvers for Vehicle Routing Problems with Faster Inference SpeedYubin Xiao, Di Wang, Boyang Li, Mingzhao Wang 等AAAI 2024 · 被引用 34 次
- Equity-Transformer: Solving NP-Hard Min-Max Routing Problems as Sequential Generation with Equity ContextJiwoo Son, Minsu Kim, Sanghyeok Choi, Hyeonah Kim 等AAAI 2024 · 被引用 28 次
- Siamese-Discriminant Deep Reinforcement Learning for Solving Jigsaw Puzzles with Large Eroded GapsXingke Song, Jiahuan Jin, Chenglin Yao, Shihe Wang 等AAAI 2023 · 被引用 24 次
- PARCO: Parallel AutoRegressive Models for Multi-Agent Combinatorial OptimizationFederico Berto, Chuanbo Hua, Laurin Luttmann, Jiwoo Son 等NeurIPS 2025 · 被引用 14 次
它引用的顶会 Paper2
相关 Paper
- Multi-Agent Pointer Transformer: Seq-to-Seq Reinforcement Learning for Multi-Vehicle Dynamic Pickup-Delivery ProblemsZengyu Zou, Jingyuan Wang, Yixuan Huang, Junjie WuAAAI 2026
- A Hierarchical Reinforcement Learning Based Optimization Framework for Large-scale Dynamic Pickup and Delivery ProblemsYi Ma, Xiaotian Hao, Jianye Hao, Jiawen Lu 等NeurIPS 2021 · 被引用 100 次
- Rethinking Light Decoder-based Solvers for Vehicle Routing ProblemsZiwei Huang, Jianan Zhou, Zhiguang Cao, Yixin XuICLR 2025
- MTL-KD: Multi-Task Learning Via Knowledge Distillation for Generalizable Neural Vehicle Routing SolverYuepeng Zheng, Fu Luo, Zhenkun Wang, Yaoxin Wu 等NeurIPS 2025 · 被引用 13 次
- Chain-of-Context Learning: Dynamic Constraint Understanding for Multi-Task VRPsShuangchun Gui, Suyu Liu, Xuehe Wang, Zhiguang CaoICLR 2026 · 被引用 3 次
