UTFix: Change Aware Unit Test Repairing using LLM
Shanto Rahman, Sachit Kuhar, Berk Çirisci, Pranav Garg, Shiqi Wang, Xiaofei Ma, Anoop Deoras, Baishakhi Ray
摘要
USA XIAOFEI MA, ANOOP DEORAS, and BAISHAKHI RAY, AWS, USA Software updates, including bug repair and feature additions, are frequent in modern applications but they often leave test suites outdated, resulting in undetected bugs and increased chances of system failures. A recent study by Meta revealed that 14%-22% of software failures stem from outdated tests that fail to reflect changes in the codebase. This highlights the need to keep tests in sync with code changes to ensure software reliability.
In this paper, we present UTFix, a novel approach for repairing unit tests when their corresponding focal methods undergo changes. UTFix addresses two critical issues: assertion failure and reduced code coverage caused by changes in the focal method. Our approach leverages language models to repair unit tests by providing contextual information such as static code slices, dynamic code slices, and failure messages. We evaluate UTFix on our generated synthetic benchmarks (Tool-Bench), and real-world benchmarks. Tool-Bench includes diverse changes from popular open-source Python GitHub projects, where UTFix successfully repaired 89.2% of assertion failures and achieved 100% code coverage for 96 tests out of 369 tests. On the real-world benchmarks, UTFix repairs 60% of assertion failures while achieving 100% code coverage for 19 out of 30 unit tests. To the best of our knowledge, this is the first comprehensive study focused on unit test in evolving Python projects. Our contributions include the development of UTFix, the creation of Tool-Bench and real-world benchmarks, and the demonstration of the effectiveness of LLM-based methods in addressing unit test failures due to software evolution.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- When Specifications Meet Reality: Uncovering API Inconsistencies in Ethereum InfrastructureJie Ma, Ningyu He, Jinwen Xi, Mingzhe Xing 等OOPSLA 2026
- Diffploit: Facilitating Cross-Version Exploit Migration for Open Source Library VulnerabilitiesZirui Chen, Zhipeng Xue, Jiayuan Zhou, Xing Hu 等ICSE 2026
- Direct Manipulation and Natural Language Programming, Together at Last?Parker Ziegler, David Minh-Duy Cao, Justin Lubin, Sarah E. ChasinsOOPSLA 2026
它引用的顶会 Paper8
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- TOGA: A Neural Method for Test Oracle GenerationElizabeth Dinella, Gabriel Ryan, Todd Mytkowicz, Shuvendu K. LahiriICSE 2022 · 被引用 92 次
- Code-Aware Prompting: A Study of Coverage-Guided Test Generation in Regression Setting using LLMGabriel Ryan, Siddhartha Jain, Mingyue Shang, Shiqi Wang 等FSE 2024 · 被引用 68 次
- Using pre-trained language models to resolve textual and semantic merge conflicts (experience paper)Jialu Zhang, Todd Mytkowicz, Mike Kaufman, Ruzica Piskac 等ISSTA 2022 · 被引用 30 次
- HITS: High-coverage LLM-based Unit Test Generation via Method SlicingZejun Wang, Kaibo Liu, Ge Li, Zhi JinASE 2024 · 被引用 29 次
相关 Paper
- Unit Test Update through LLM-Driven Context Collection and Error-Type-Aware RefinementYuanhe Zhang, Zhiquan Yang, Shengyi Pan, Zhongxin LiuASE 2025 · 被引用 3 次
- MuMuTestUp: Mutation-Based Multi-agent Test Case UpdateDawei Tian, Jiakun Liu, Yun Peng, Yichen Zhang 等ISSTA 2026
- Comprehend, Imitate, and then Update: Unleashing the Power of LLMs in Test Suite EvolutionTangzhi Xu, Jianhan Liu, Yuan Yao, Cong Li 等ASE 2025 · 被引用 1 次
- On the Evaluation of Large Language Models in Unit Test Evolution (Experience Paper)Weichang Liu, Junwei Zhang, Yuqing Niu, Bo ZhouISSTA 2026
- CodeSync: Synchronizing Large Language Models with Dynamic Code Evolution at ScaleChenlong Wang, Zhaoyang Chu, Zhengxiang Cheng, Xuyi Yang 等ICML 2025
