Refining Hybrid Genetic Search for CVRP via Reinforcement Learning-Finetuned LLM
Rongjie Zhu, Cong Zhang, Zhiguang Cao
Abstract
While large language models (LLMs) are emerging as automated heuristic designers for solving vehicle routing problems (VRPs), state-of-the-art approaches predominantly rely on massive, general-purpose models like GPT-4. This work challenges this paradigm by demonstrating that smaller, specialized LLMs, when finely tuned, can generate components that surpass expert-designed heuristics within advanced solvers. We introduce RFTHGS, a novel Reinforcement learning (RL) framework for Fine-Tuning a small LLM to produce high-performance crossover operators for the Hybrid Genetic Search (HGS) solver to solve the capacitated vehicle routing problem (CVRP). Our methods utilizes a multi-tiered, curriculum-based reward function that progressively guides the LLM to first produce compilable code, then executable operators, and finally, components that exceed human expert-designed ones. Additionally, we introduce an operator caching mechanism to work in conjunction with the reward function, discouraging plagiarism and promoting diversity during training. Experimental results demonstrate that our fine-tuned LLM generates crossover operators which significantly outperform those designed by human experts in HGS. This performance advantage is consistent, holding from small-scale instances and generalizing to large-scale problems of up to 1000 nodes. Furthermore, RFTHGS surpasses leading neurocombinatorial baselines, prompt-based methods, and commercial LLMs, including GPT-4o and GPT-4o-mini.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b4d9e396-51b2-4a84-a805-287ab9667286Builds on15
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Large Language Models as OptimizersChengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu et al.ICLR 2024 · 817 citations
- POMO: Policy Optimization with Multiple Optima for Reinforcement LearningYeong-Dae Kwon, Jinho Choo, Byoungjip Kim, Iljoo Yoon et al.NeurIPS 2020 · 731 citations
Related papers
- Large Language Models as End-to-end Combinatorial Optimization SolversXia Jiang, Yaoxin Wu, Minshuo Li, Zhiguang Cao et al.NeurIPS 2025 · 37 citations
- An Agentic Framework with LLMs for Solving Complex Vehicle Routing ProblemsNi Zhang, Zhiguang Cao, Jianan Zhou, Cong Zhang et al.ICLR 2026 · 8 citations
- DRoC: Elevating Large Language Models for Complex Vehicle Routing via Decomposed Retrieval of ConstraintsXia Jiang, Yaoxin Wu, Chenhao Zhang, Yingqian ZhangICLR 2025
- Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement LearningSimon Zhai, Hao Bai, Zipeng Lin, Jiayi Pan et al.NeurIPS 2024 · 214 citations
- A Learning-based Iterative Method for Solving Vehicle Routing ProblemsHao Lu, Xingwen Zhang, Shuang YangICLR 2020 · 270 citations
