Can Old Tests Do New Tricks for Resolving SWE Issues?
Yang Chen, Toufique Ahmed, Reyhaneh Jabbarvand, Martin Hirzel
摘要
MARTIN HIRZEL, IBM, USA Test suites in real-world projects are often large and achieve high code coverage, yet they remain insufficient for detecting all bugs. The abundance of unresolved issues in open-source project trackers highlights this gap. While regression tests are typically designed to ensure past functionality is preserved in the new version, they can also serve a complementary purpose: debugging the current version. Specifically, regression tests can (1) enhance the generation of reproduction tests for newly reported issues, and (2) validate that patches do not regress existing functionality. We present TestPrune, a fully automated technique that leverages issue tracker reports and strategically reuses regression tests for both bug reproduction and patch validation.
A key contribution of TestPrune is its ability to automatically minimize the regression suite to a small, highly relevant subset of tests. Due to the predominance of LLM-based debugging techniques, this minimization is essential as large test suites exceed context limits, introduce noise, and inflate inference costs. TestPrune can be plugged into any agentic bug repair pipeline and orthogonally improve overall performance. As a proof of concept, we show that TestPrune leads to a 6.2% -9.0% relative increase in issue reproduction rate within the Otter framework and a 8.0% -12.9% relative increase in issue resolution rate within Agentless, SWE-Agent, and Trae agent on SWE-Bench Lite and SWE-Bench Verified benchmarks. Compared to the benefits, the model API cost overhead of TestPrune is minimal, at 0.05 per SWE-Bench instance using GPT-4o and Claude-3.7-Sonnet models, respectively.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper19
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
- SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software EvolutionYuxiang Wei, Olivier Duchenne, Jade Copet, Quentin Carbonneaux 等NeurIPS 2025 · 被引用 291 次
- SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code AgentsNiels Mündler, Mark Niklas Müller, Jingxuan He, Martin T. VechevNeurIPS 2024 · 被引用 172 次
- AutoCodeRover: Autonomous Program ImprovementYuntong Zhang, Haifeng Ruan, Zhiyu Fan, Abhik RoychoudhuryISSTA 2024 · 被引用 96 次
- Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of Large Language ModelsMosh Levy, Alon Jacoby, Yoav GoldbergACL 2024 · 被引用 77 次
相关 Paper
- RegMiner: towards constructing a large regression dataset from code evolution historyXuezhi Song, Yun Lin, Siang Hwee Ng, Yijian Wu 等ISSTA 2022 · 被引用 11 次
- Issue2Test: Generating Reproducing Test Cases from Issue ReportsNoor Nashid, Islem Bouzenia, Michael Pradel, Ali MesbahICSE 2026 · 被引用 1 次
- Testora: Using Natural Language Intent to Detect Behavioral RegressionsMichael PradelICSE 2026 · 被引用 1 次
- Otter: Generating Tests from Issues to Validate SWE PatchesToufique Ahmed, Jatin Ganhotra, Rangeet Pan, Avraham Shinnar 等ICML 2025
- iCoRe: An Iterative Correlation-Aware Retriever for Bug Reproduction Test GenerationJunyi Wang, Jialun Cao, Zhongxin LiuFSE 2026
