Lune

NeurIPS2025顶会

Co-PatcheR: Collaborative Software Patching with Component-specific Small Reasoning Models

Yuheng Tang, Hongwei Li, Kaijie Zhu, Michael Yang, Yangruibo Ding, Wenbo Guo

2025年份
2被引次数

摘要

Motivated by the success of general-purpose large language models (LLMs) in software patching, recent works started to train specialized patching models. Most works trained one model to handle the end-to-end patching pipeline (including issue localization, patch generation, and patch validation). However, it is hard for a small model to handle all tasks, as different sub-tasks have different workflows and require different expertise. As such, by using a 70 billion model, SOTA methods can only reach up to 41% resolved rate on SWE-bench-Verified. Motivated by the collaborative nature, we propose Co-PatcheR, the first collaborative patching system with small and specialized reasoning models for individual components.

Our key technique novelties are the specific task designs and training recipes. First, we train a model for localization and patch generation. Our localization pinpoints the suspicious lines through a two-step procedure, and our generation combines patch generation and critique. We then propose a hybrid patch validation that includes two models for crafting issue-reproducing test cases with and without assertions and judging patch correctness, followed by a majority vote-based patch selection. Through extensive evaluation, we show that Co-PatcheR achieves 46% resolved rate on SWE-bench-Verified with only 3×14B models This makes Co-PatcheR the best patcher with specialized models, requiring the least training resources and the smallest models. We conduct a comprehensive ablation study to validate our recipes, as well as our choice of training data number, model size, and testing-phase scaling strategy.

Through extensive experiments, we first show that when using only 3×14B models, Co-PatcheR can achieve a 46% resolved rate on SWE-bench-Verified with 60 patch candidates. Compared to SWE-RL, Co-PatcheR achieves a high resolved rate with 40% fewer parameters and 88% fewer samples. Besides, Co-PatcheR only needs to run one 14B model at a time, which is much more efficient than SOTA methods during the testing phase. Furthermore, with our specific reasoning data construction method, Co-PatcheR only requires 6K data for training, which is much more efficient than SOTA methods that use at least 30K samples. We then conduct a comprehensive ablation study for each model to validate its task design and training recipe. Finally, we validate the necessity of testing-phase reasoning, our choice of data number and model size, through more ablation studies.

Contributions. We propose Co-PatcheR, the first collaborative patching system with componentspecific reasoning models. Co-PatcheR is the most data-and parameter-efficient patcher that offers greater effectiveness, efficiency, and modularity than existing patchers with specialized models. Co-PatcheR ranks among the top-10 open-source systems on SWE-bench-Verified, outperforming all patchers with open-source models. We propose specific training recipes for each model and obtain the following new findings that are unique to patching:

• Using one model for localization and generation performs similarly to using separate models.

• Multiple models for PoC generation provide necessary diversity that a single model cannot achieve.

• Critique is important for generation, and multi-source data is important for validation.

• Simply increasing data or model size is not always helpful; data scale should match model size.

• Rejection sampling-based data filtering helps all components; but rationalization does not.

2 Existing Works and Limitations LLM-based patching agent. There are several works on designing a patching agent using generalpurpose LLMs [25,7,6,27,2,20, 8, 14,13,56,59,5,42,35]. Some agents achieve remarkable performance on the SWE-bench benchmark [22], a benchmark for real-world GitHub issues written in Python. The top-ranked open-source agents are OpenHands [48], Agentless [53], and PatchPilot [24]. Here, Agentless and PatchPilot follow a pre-defined workflow, where PatchPilot introduces a number

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext eaefbe28-bfd5-436b-afd8-cde55ea84b77

它引用的顶会 Paper21

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖