A Hierarchical Reinforcement Learning Based Optimization Framework for Large-scale Dynamic Pickup and Delivery Problems
Yi Ma, Xiaotian Hao, Jianye Hao, Jiawen Lu, Xing Liu, Xialiang Tong, Mingxuan Yuan, Zhigang Li, Jie Tang, Zhaopeng Meng
Abstract
The Dynamic Pickup and Delivery Problem (DPDP) is an essential problem in the logistics domain, which is NP-hard. The objective is to dynamically schedule vehicles among multiple sites to serve the online generated orders such that the overall transportation cost could be minimized. The critical challenge of DPDP is the orders are not known a priori, i.e., the orders are dynamically generated in real-time. To address this problem, existing methods partition the overall DPDP into fixed-size sub-problems by caching online generated orders and solve each sub-problem, or on this basis to utilize the predicted future orders to optimize each sub-problem further. However, the solution quality and efficiency of these methods are unsatisfactory, especially when the problem scale is very large. In this paper, we propose a novel hierarchical optimization framework to better solve large-scale DPDPs. Specifically, we design an upper-level agent to dynamically partition the DPDP into a series of sub-problems with different scales to optimize vehicles routes towards globally better solutions. Besides, a lower-level agent is designed to efficiently solve each sub-problem by incorporating the strengths of classical operational research-based methods with reinforcement learning-based policies. To verify the effectiveness of the proposed framework, real historical data is collected from the order dispatching system of Huawei Supply Chain Business Unit and used to build a functional simulator. Extensive offline simulation and online testing conducted on the industrial order dispatching system justify the superior performance of our framework over existing baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4ab72950-b88f-4adf-9ff4-44a35b4b34aeCited by top-tier papers11
- H-TSP: Hierarchically Solving the Large-Scale Traveling Salesman ProblemXuanhao Pan, Yan Jin, Yuandong Ding, Mingxiao Feng et al.AAAI 2023 · 85 citations
- Hierarchical Diffusion for Offline Decision MakingWenhao Li, Xiangfeng Wang, Bo Jin, Hongyuan ZhaICML 2023 · 80 citations
- Equity-Transformer: Solving NP-Hard Min-Max Routing Problems as Sequential Generation with Equity ContextJiwoo Son, Minsu Kim, Sanghyeok Choi, Hyeonah Kim et al.AAAI 2024 · 28 citations
- GAT-MF: Graph Attention Mean Field for Very Large Scale Multi-Agent Reinforcement LearningQianyue Hao, Wenzhen Huang, Tao Feng, Jian Yuan et al.KDD 2023 · 18 citations
- Efficient Planning with Latent DiffusionWenhao LiICLR 2024 · 15 citations
Builds on1
Related papers
- MAPDP: Cooperative Multi-Agent Reinforcement Learning to Solve Pickup and Delivery ProblemsZefang Zong, Meng Zheng, Yong Li, Depeng JinAAAI 2022 · 66 citations
- Multi-Agent Pointer Transformer: Seq-to-Seq Reinforcement Learning for Multi-Vehicle Dynamic Pickup-Delivery ProblemsZengyu Zou, Jingyuan Wang, Yixuan Huang, Junjie WuAAAI 2026
- NondBREM: Nondeterministic Offline Reinforcement Learning for Large-Scale Order DispatchingHongbo Zhang, Guang Wang, Xu Wang, Zhengyang Zhou et al.AAAI 2024 · 9 citations
- Rolling Horizon Based Temporal Decomposition for the Offline Pickup and Delivery Problem with Time WindowsYoungseo Kim, Danushka Edirimanna, Michael Wilbur, Philip Pugliese et al.AAAI 2023 · 11 citations
- Online Trichromatic Pickup and Delivery Scheduling in Spatial CrowdsourcingBolong Zheng, Chenze Huang, Christian S. Jensen, Lu Chen et al.ICDE 2020 · 34 citations
