Triple-BERT: Do We Really Need MARL for Order Dispatch on Ride-Sharing Platforms?
Zijian Zhao, Sen Li
摘要
On-demand ride-sharing platforms, such as Uber and Lyft, face the intricate realtime challenge of bundling and matching passengers-each with distinct origins and destinations-to available vehicles, all while navigating significant system uncertainties. Due to the extensive observation space arising from the large number of drivers and orders, order dispatching, though fundamentally a centralized task, is often addressed using Multi-Agent Reinforcement Learning (MARL). However, independent MARL methods fail to capture global information and exhibit poor cooperation among workers, while Centralized Training Decentralized Execution (CTDE) MARL methods suffer from the Curse of Dimensionality (CoD). To overcome these challenges, we propose Triple-BERT, a centralized framework designed specifically for large-scale order dispatching on ride-sharing platforms based on Single Agent Reinforcement Learning (SARL). Built on a variant TD3, our approach addresses the vast action space through an action decomposition strategy that breaks down the joint action probability into individual driver action probabilities. To handle the extensive observation space, we introduce a novel BERT-based network, where parameter reuse mitigates parameter growth as the number of drivers and orders increases, and the attention mechanism effectively captures the complex relationships among the large pool of driver and orders. We validate our method using a real-world ride-hailing dataset from Manhattan. Triple-BERT achieves approximately an 11.95% improvement over current state-ofthe-art methods, with a 4.26% increase in served orders and a 22.25% reduction in pickup times. Our code, trained model parameters, and processed data are publicly available at the repository https://github.com/RS2002/Triple-BERT .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Value-Decomposition Multi-Agent Actor-CriticsJianyu Su, Stephen C. Adams, Peter A. BelingAAAI 2021 · 被引用 140 次
- Shapley Counterfactual Credits for Multi-Agent Reinforcement LearningJiahui Li, Kun Kuang, Baoxiang Wang, Furui Liu 等KDD 2021 · 被引用 49 次
- DyPS: Dynamic Parameter Sharing in Multi-Agent Reinforcement Learning for Spatio-Temporal Resource AllocationJingwei Wang, Qianyue Hao, Wenzhen Huang, Xiaochen Fan 等KDD 2024 · 被引用 7 次
- Multi Agent Reinforcement Learning for Sequential Satellite Assignment ProblemsJoshua Holder, Natasha Jaques, Mehran MesbahiAAAI 2025 · 被引用 5 次
相关 Paper
- NondBREM: Nondeterministic Offline Reinforcement Learning for Large-Scale Order DispatchingHongbo Zhang, Guang Wang, Xu Wang, Zhengyang Zhou 等AAAI 2024 · 被引用 9 次
- CoopRide: Cooperate All Grids in City-Scale Ride-Hailing Dispatching with Multi-Agent Reinforcement LearningJingwei Wang, Qianyue Hao, Wenzhen Huang, Xiaochen Fan 等KDD 2025 · 被引用 4 次
- A Queueing-Theoretic Framework for Vehicle Dispatching in Dynamic Car-HailingPeng Cheng, Jiabao Jin, Lei Chen, Xuemin Lin 等VLDB 2021 · 被引用 18 次
- H-TSP: Hierarchically Solving the Large-Scale Traveling Salesman ProblemXuanhao Pan, Yan Jin, Yuandong Ding, Mingxiao Feng 等AAAI 2023 · 被引用 85 次
- GARLIC: GPT-Augmented Reinforcement Learning with Intelligent Control for Vehicle DispatchingXiao Han, Zijian Zhang, Xiangyu Zhao, Yuanshao Zhu 等AAAI 2025 · 被引用 10 次
