Multi-Agent Pointer Transformer: Seq-to-Seq Reinforcement Learning for Multi-Vehicle Dynamic Pickup-Delivery Problems
Zengyu Zou, Jingyuan Wang, Yixuan Huang, Junjie Wu
Abstract
This paper addresses the cooperative Multi-Vehicle Dynamic Pickup and Delivery Problem with Stochastic Requests (MVDPDPSR) and proposes an end-to-end centralized decision-making framework based on sequence-to-sequence, named Multi-Agent Pointer Transformer (MAPT). MVD-PDPSR is an extension of the vehicle routing problem and a spatio-temporal system optimization problem, widely applied in scenarios such as on-demand delivery. Classical operations research methods face bottlenecks in computational complexity and time efficiency when handling large-scale dynamic problems. Although existing reinforcement learning methods have achieved some progress, they still encounter several challenges: 1) Independent decoding across multiple vehicles fails to model joint action distributions; 2) The feature extraction network struggles to capture inter-entity relationships; 3) The joint action space is exponentially large. To address these issues, we designed the MAPT framework, which employs a Transformer Encoder to extract entity representations, combines a Transformer Decoder with a Pointer Network to generate joint action sequences in an AutoRegressive manner, and introduces a Relation-Aware Attention module to capture inter-entity relationships. Additionally, we guide the model's decision-making using informative priors to facilitate effective exploration. Experiments on 8 datasets demonstrate that MAPT significantly outperforms existing baseline methods in terms of performance and exhibits substantial computational time advantages compared to classical operations research methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f6e4f6dd-c6ff-4f24-a0a1-e46472dd9318Builds on9
- Multi-Agent Reinforcement Learning is a Sequence Modeling ProblemMuning Wen, Jakub Grudzien Kuba, Runji Lin, Weinan Zhang et al.NeurIPS 2022 · 408 citations
- Learning Effective Road Network Representation with Hierarchical Graph Neural NetworksNing Wu, Wayne Xin Zhao, Jingyuan Wang, Dayan PanKDD 2020 · 109 citations
- Self-supervised Trajectory Representation Learning with Temporal Regularities and Travel SemanticsJiawei Jiang, Dayan Pan, Houxing Ren, Xiaohan Jiang et al.ICDE 2023 · 101 citations
- Continuous Trajectory Generation Based on Two-Stage GANWenjun Jiang, Wayne Xin Zhao, Jingyuan Wang, Jiawei JiangAAAI 2023 · 77 citations
- MAPDP: Cooperative Multi-Agent Reinforcement Learning to Solve Pickup and Delivery ProblemsZefang Zong, Meng Zheng, Yong Li, Depeng JinAAAI 2022 · 66 citations
Related papers
- Pointerformer: Deep Reinforced Multi-Pointer Transformer for the Traveling Salesman ProblemYan Jin, Yuandong Ding, Xuanhao Pan, Kun He et al.AAAI 2023 · 79 citations
- Equity-Transformer: Solving NP-Hard Min-Max Routing Problems as Sequential Generation with Equity ContextJiwoo Son, Minsu Kim, Sanghyeok Choi, Hyeonah Kim et al.AAAI 2024 · 28 citations
- A Hierarchical Reinforcement Learning Based Optimization Framework for Large-scale Dynamic Pickup and Delivery ProblemsYi Ma, Xiaotian Hao, Jianye Hao, Jiawen Lu et al.NeurIPS 2021 · 100 citations
- Continuous Spatiotemporal TransformerAntonio Henrique de Oliveira Fonseca, Emanuele Zappala, Josue Ortega Caro, David van DijkICML 2023 · 2 citations
- Decoding Global Preferences: Temporal and Cooperative Dependency Modeling in Multi-Agent Preference-Based Reinforcement LearningTianchen Zhu, Yue Qiu, Haoyi Zhou, Jianxin LiAAAI 2024 · 9 citations
