Template-Guided Program Repair in the Era of Large Language Models
Kai Huang, Jian Zhang, Xiangxin Meng, Yang Liu
Abstract
Recent advancements in automated program repair (APR) have been significantly driven by the application of Large Language Models (LLMs). In particular, the integration of LLMs with traditional template-based repair methods has demonstrated effective outcomes. Despite this, the synergy between the strengths of traditional methods and LLMs remains underexploited. This oversight originates from the indiscriminate use of templates and their insufficient coverage. Also, using small-scale LLMs within the zero-shot learning context proves to be suboptimal. To alleviate the limitations, we propose NTR (Neural Template Repair), a two-stage repair framework including template selection and patch generation, both of which are under the fine-tuning paradigm. In the template selection phase, we formulate it as a multiclass classification problem and fine-tune million-level LLMs for better selecting possible templates. During the patch generation phase, we leverage the chosen templates as probable directions (e.g., ‘Mutate Conditional Expression’) to guide the fine-tuning process of LLMs at the billion-level scale for precise patch creation. Moreover, we incorporate a unique template to signify the absence of a suitable template and employ a probability-based prioritization of templates, thereby optimizing patch generation. This framework not only effectively addresses template mismatch issues, but also enables the billion-level LLMs to explore the patch space more efficiently, despite the GPU memory constraints. We evaluate NTR with different foundational models on Defects4J V1.2 and HumanEval-Java, the framework consistently demonstrates significant effectiveness. When utilizing StarCoder as the foundational model for patch generation, NTR fixes 128 and 129 bugs in Defects4J and HumanEval, outperforming the best baseline APR tool by 14 and 59 bugs. With the larger CodeLlama model, the fixed bugs rise to 139 and 136, respectively, exceeding the baseline by 25 and 66 bugs. Notably, the performance stems not only from the foundational models but also benefits greatly from our NTR framework. Specifically, NTR's implementation with StarCoder and CodeLlama leads to 22 and 23 additional fixes, which is beyond what the models achieve on their own. This emphasizes the success of our new perspective on utilizing templates to unlock the bug-fixing potential of LLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b8a37aec-86cc-49c2-89cf-1f62b9e89ae9Cited by top-tier papers6
- Seeing is Fixing: Cross-Modal Reasoning with Multimodal LLMs for Visual Software Issue RepairKai Huang, Jian Zhang, Xiaofei Xie, Chunyang ChenASE 2025 · 5 citations
- TLR: Codebase-Level C Memory Management Error Repair with Large Language ModelsXiao Cheng, Zhihao Guo, Huan Huo, Yulei SuiFSE 2026 · 1 citation
- Vul-R2: A Reasoning LLM for Automated Vulnerability RepairXin-Cheng Wen, Zirui Lin, Yijun Yang, Cuiyun Gao et al.ASE 2025 · 1 citation
- SoK: Automated Vulnerability Repair: Methods, Tools, and AssessmentsYiwei Hu, Zhen Li, Kedie Shu, Shenghua Guan et al.USENIX Security 2025
- VulKey: Automated Vulnerability Repair Guided by Domain-Specific Repair PatternsJia Li, Zhuangbin Chen, Yuxin Su, Michael R. LyuFSE 2026
Builds on32
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 1,224 citations
- CoCoNuT: combining context-aware neural translation models using ensemble for program repairThibaud Lutellier, Hung Viet Pham, Lawrence Pang, Yitong Li et al.ISSTA 2020 · 325 citations
- CURE: Code-Aware Neural Machine Translation for Automatic Program RepairNan Jiang, Thibaud Lutellier, Lin TanICSE 2021 · 267 citations
Related papers
- Automated Program Repair in the Era of Large Pre-trained Language ModelsChunqiu Steven Xia, Yuxiang Wei, Lingming ZhangICSE 2023 · 321 citations
- Demystifying Memorization in LLM-Based Program Repair via a General Hypothesis Testing FrameworkJiaolong Kong, Xiaofei Xie, Shangqing LiuFSE 2025 · 5 citations
- ThinkRepair: Self-Directed Automated Program RepairXin Yin, Chao Ni, Shaohua Wang, Zhenhao Li et al.ISSTA 2024 · 37 citations
- Template-based Neural Program RepairXiangxin Meng, Xu Wang, Hongyu Zhang, Hailong Sun et al.ICSE 2023 · 29 citations
- The Plastic Surgery Hypothesis in the Era of Large Language ModelsChunqiu Steven Xia, Yifeng Ding, Lingming ZhangASE 2023 · 26 citations
