Refining Hybrid Genetic Search for CVRP via Reinforcement Learning-Finetuned LLM
Rongjie Zhu, Cong Zhang, Zhiguang Cao
摘要
While large language models (LLMs) are emerging as automated heuristic designers for solving vehicle routing problems (VRPs), state-of-the-art approaches predominantly rely on massive, general-purpose models like GPT-4. This work challenges this paradigm by demonstrating that smaller, specialized LLMs, when finely tuned, can generate components that surpass expert-designed heuristics within advanced solvers. We introduce RFTHGS, a novel Reinforcement learning (RL) framework for Fine-Tuning a small LLM to produce high-performance crossover operators for the Hybrid Genetic Search (HGS) solver to solve the capacitated vehicle routing problem (CVRP). Our methods utilizes a multi-tiered, curriculum-based reward function that progressively guides the LLM to first produce compilable code, then executable operators, and finally, components that exceed human expert-designed ones. Additionally, we introduce an operator caching mechanism to work in conjunction with the reward function, discouraging plagiarism and promoting diversity during training. Experimental results demonstrate that our fine-tuned LLM generates crossover operators which significantly outperform those designed by human experts in HGS. This performance advantage is consistent, holding from small-scale instances and generalizing to large-scale problems of up to 1000 nodes. Furthermore, RFTHGS surpasses leading neurocombinatorial baselines, prompt-based methods, and commercial LLMs, including GPT-4o and GPT-4o-mini.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Large Language Models as OptimizersChengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu 等ICLR 2024 · 被引用 817 次
- POMO: Policy Optimization with Multiple Optima for Reinforcement LearningYeong-Dae Kwon, Jinho Choo, Byoungjip Kim, Iljoo Yoon 等NeurIPS 2020 · 被引用 731 次
相关 Paper
- Large Language Models as End-to-end Combinatorial Optimization SolversXia Jiang, Yaoxin Wu, Minshuo Li, Zhiguang Cao 等NeurIPS 2025 · 被引用 37 次
- An Agentic Framework with LLMs for Solving Complex Vehicle Routing ProblemsNi Zhang, Zhiguang Cao, Jianan Zhou, Cong Zhang 等ICLR 2026 · 被引用 8 次
- DRoC: Elevating Large Language Models for Complex Vehicle Routing via Decomposed Retrieval of ConstraintsXia Jiang, Yaoxin Wu, Chenhao Zhang, Yingqian ZhangICLR 2025
- Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement LearningSimon Zhai, Hao Bai, Zipeng Lin, Jiayi Pan 等NeurIPS 2024 · 被引用 214 次
- A Learning-based Iterative Method for Solving Vehicle Routing ProblemsHao Lu, Xingwen Zhang, Shuang YangICLR 2020 · 被引用 270 次
