Multi-Agent Pointer Transformer: Seq-to-Seq Reinforcement Learning for Multi-Vehicle Dynamic Pickup-Delivery Problems
Zengyu Zou, Jingyuan Wang, Yixuan Huang, Junjie Wu
摘要
This paper addresses the cooperative Multi-Vehicle Dynamic Pickup and Delivery Problem with Stochastic Requests (MVDPDPSR) and proposes an end-to-end centralized decision-making framework based on sequence-to-sequence, named Multi-Agent Pointer Transformer (MAPT). MVD-PDPSR is an extension of the vehicle routing problem and a spatio-temporal system optimization problem, widely applied in scenarios such as on-demand delivery. Classical operations research methods face bottlenecks in computational complexity and time efficiency when handling large-scale dynamic problems. Although existing reinforcement learning methods have achieved some progress, they still encounter several challenges: 1) Independent decoding across multiple vehicles fails to model joint action distributions; 2) The feature extraction network struggles to capture inter-entity relationships; 3) The joint action space is exponentially large. To address these issues, we designed the MAPT framework, which employs a Transformer Encoder to extract entity representations, combines a Transformer Decoder with a Pointer Network to generate joint action sequences in an AutoRegressive manner, and introduces a Relation-Aware Attention module to capture inter-entity relationships. Additionally, we guide the model's decision-making using informative priors to facilitate effective exploration. Experiments on 8 datasets demonstrate that MAPT significantly outperforms existing baseline methods in terms of performance and exhibits substantial computational time advantages compared to classical operations research methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Multi-Agent Reinforcement Learning is a Sequence Modeling ProblemMuning Wen, Jakub Grudzien Kuba, Runji Lin, Weinan Zhang 等NeurIPS 2022 · 被引用 408 次
- Learning Effective Road Network Representation with Hierarchical Graph Neural NetworksNing Wu, Wayne Xin Zhao, Jingyuan Wang, Dayan PanKDD 2020 · 被引用 109 次
- Self-supervised Trajectory Representation Learning with Temporal Regularities and Travel SemanticsJiawei Jiang, Dayan Pan, Houxing Ren, Xiaohan Jiang 等ICDE 2023 · 被引用 101 次
- Continuous Trajectory Generation Based on Two-Stage GANWenjun Jiang, Wayne Xin Zhao, Jingyuan Wang, Jiawei JiangAAAI 2023 · 被引用 77 次
- MAPDP: Cooperative Multi-Agent Reinforcement Learning to Solve Pickup and Delivery ProblemsZefang Zong, Meng Zheng, Yong Li, Depeng JinAAAI 2022 · 被引用 66 次
相关 Paper
- Pointerformer: Deep Reinforced Multi-Pointer Transformer for the Traveling Salesman ProblemYan Jin, Yuandong Ding, Xuanhao Pan, Kun He 等AAAI 2023 · 被引用 79 次
- Equity-Transformer: Solving NP-Hard Min-Max Routing Problems as Sequential Generation with Equity ContextJiwoo Son, Minsu Kim, Sanghyeok Choi, Hyeonah Kim 等AAAI 2024 · 被引用 28 次
- A Hierarchical Reinforcement Learning Based Optimization Framework for Large-scale Dynamic Pickup and Delivery ProblemsYi Ma, Xiaotian Hao, Jianye Hao, Jiawen Lu 等NeurIPS 2021 · 被引用 100 次
- Continuous Spatiotemporal TransformerAntonio Henrique de Oliveira Fonseca, Emanuele Zappala, Josue Ortega Caro, David van DijkICML 2023 · 被引用 2 次
- Decoding Global Preferences: Temporal and Cooperative Dependency Modeling in Multi-Agent Preference-Based Reinforcement LearningTianchen Zhu, Yue Qiu, Haoyi Zhou, Jianxin LiAAAI 2024 · 被引用 9 次
