PoMtVRS: Preference-Optimized Multi-Task Vehicle Routing Solver with Preference Gating
Dian Meng, Yaoxin Wu, Yaqing Hou, Zhiguang Cao
Abstract
Multi-task vehicle routing solvers via deep reinforcement learning have attracted broad attention and achieved significant progress in handling multiple constraints. However, existing neural solvers still face critical challenges, including insufficient representation, unstable training, and inefficient exploration in large combinatorial action spaces, which often prevents performance from meeting its full potential. To address these issues, we propose PoMtVRS (Preference-Optimized Multi-Task Vehicle Routing Solver with Preference Gating), a plug-and-play framework that jointly improves decoder representations and exploration efficiency through a synergistic combination of decoder-side augmentation and preference-driven optimization. Specifically, we introduce the preference optimization objective to learn relative comparisons among candidate solutions for different routing tasks, encouraging a higher generation probability of better solutions. Meanwhile, we design a preference-gated block that adaptively modulates decoder representations via sparse gated attention and nonlinear residual refinement. Extensive experiments demonstrate that PoMtVRS elevates state-of-the-art unified neural VRP backbones, achieving leading performance in multi-task benchmarks and stronger generalization. The code is available at: https: //github.com/Regina921/PoMtVRS.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 19de7d23-3bcb-442c-b254-e5a2bf99f7c3Builds on29
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Preference Ranking Optimization for Human AlignmentFeifan Song, Bowen Yu, Minghao Li, Haiyang Yu et al.AAAI 2024 · 357 citations
- Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-FreeZihan Qiu, Zekun Wang, Bo Zheng, Zeyu Huang et al.NeurIPS 2025 · 336 citations
- A Learning-based Iterative Method for Solving Vehicle Routing ProblemsHao Lu, Xingwen Zhang, Shuang YangICLR 2020 · 270 citations
Related papers
- MVMoE: Multi-Task Vehicle Routing Solver with Mixture-of-ExpertsJianan Zhou, Zhiguang Cao, Yaoxin Wu, Wen Song et al.ICML 2024 · 74 citations
- URS: A Unified Neural Routing Solver for Cross-Problem Zero-Shot GeneralizationChangliang Zhou, Canhong Yu, Shunyu Yao, Xi Lin et al.ICML 2026 · 12 citations
- USPR: Learning a Unified Solver for Profiled RoutingChuanbo Hua, Federico Berto, Zhikai Zhao, Jiwoo Son et al.AAAI 2026 · 2 citations
- A Mixed-Curvature based Pre-training Paradigm for Multi-Task Vehicle Routing SolverSuyu Liu, Zhiguang Cao, Shanshan Feng, Yew-Soon OngICML 2025
- MTL-KD: Multi-Task Learning Via Knowledge Distillation for Generalizable Neural Vehicle Routing SolverYuepeng Zheng, Fu Luo, Zhenkun Wang, Yaoxin Wu et al.NeurIPS 2025 · 13 citations
