NondBREM: Nondeterministic Offline Reinforcement Learning for Large-Scale Order Dispatching
Hongbo Zhang, Guang Wang, Xu Wang, Zhengyang Zhou, Chen Zhang, Zheng Dong, Yang Wang
摘要
One of the most important tasks in ride-hailing is order dispatching, i.e., assigning unserved orders to available drivers. Recent order dispatching has achieved a significant improvement due to the advance of reinforcement learning, which has been approved to be able to effectively address sequential decision-making problems like order dispatching. However, most existing reinforcement learning methods require agents to learn the optimal policy by interacting with environments online, which is challenging or impractical for real-world deployment due to high costs or safety concerns. For example, due to the spatiotemporally unbalanced supply and demand, online reinforcement learning-based order dispatching may significantly impact the revenue of the ride-hailing platform and passenger experience during the policy learning period. Hence, in this work, we develop an offline deep reinforcement learning framework called NondBREM for large-scale order dispatching, which learns policy from only the accumulated logged data to avoid costly and unsafe interactions with the environment. In NondBREM, a Nondeterministic Batch-Constrained Q-learning (NondBCQ) module is developed to reduce the algorithm extrapolation error and a Random Ensemble Mixture (REM) module that integrates multiple value networks with multi-head networks is utilized to improve the model generalization and robustness. Extensive experiments on large-scale real-world ride-hailing datasets show the superiority of our design.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper3
- An Optimistic Perspective on Offline Reinforcement LearningRishabh Agarwal, Dale Schuurmans, Mohammad NorouziICML 2020 · 被引用 568 次
- A General Offline Reinforcement Learning Framework for Interactive RecommendationTeng Xiao, Donglin WangAAAI 2021 · 被引用 82 次
- Leveraging Factored Action Spaces for Efficient Offline Reinforcement Learning in HealthcareShengpu Tang, Maggie Makar, Michael W. Sjoding, Finale Doshi-Velez 等NeurIPS 2022 · 被引用 63 次
相关 Paper
- Triple-BERT: Do We Really Need MARL for Order Dispatch on Ride-Sharing Platforms?Zijian Zhao, Sen LiICLR 2026 · 被引用 4 次
- CoopRide: Cooperate All Grids in City-Scale Ride-Hailing Dispatching with Multi-Agent Reinforcement LearningJingwei Wang, Qianyue Hao, Wenzhen Huang, Xiaochen Fan 等KDD 2025 · 被引用 4 次
- i-Rebalance: Personalized Vehicle Repositioning for Supply Demand BalanceHaoyang Chen, Peiyan Sun, Qiyuan Song, Wanyuan Wang 等AAAI 2024 · 被引用 12 次
- A Hierarchical Reinforcement Learning Based Optimization Framework for Large-scale Dynamic Pickup and Delivery ProblemsYi Ma, Xiaotian Hao, Jianye Hao, Jiawen Lu 等NeurIPS 2021 · 被引用 100 次
- Multi-Objective Order Dispatch for Urban Crowd Sensing with For-Hire VehiclesJiahui Sun, Haiming Jin, Rong Ding, Guiyun Fan 等INFOCOM 2023 · 被引用 6 次
