CausalRepair: Bridging the Causality Gap in Large Language Model-Based Automated Program Repair via Dual-Slicing
Linhao Wu, Yizhou Chen, Zhen Yang, Pengyu Xue, Dan Hao
Abstract
Automated Program Repair (APR) aims to automatically fix buggy programs. In recent years, with the rapid advancement of Large Language Models (LLMs), LLM-based APR techniques have achieved significant progress. Despite their potential, the effectiveness of LLMs relies heavily on the quality of the provided repair context. However, existing LLM-based APR approaches suffer from a causality gap when constructing such contexts. Specifically, on the test side, existing methods struggle with test context ambiguity arising from noise interference or dependency absence; meanwhile, on the source side, existing retrieval-augmented methods primarily rely on static analysis and inevitably introduce static over-approximation, resulting in contexts filled with unexecuted code and noise. Consequently, these contexts mislead LLMs, hindering them from identifying the true root cause and leading to incorrect fixes.
To bridge this gap, we introduce the concept of minimal causal context, defined as the essential set of dependencies required to explain a specific failure. Based on this, we propose CausalRepair, a novel conversation-driven APR framework that instantiates this concept through a synergistic dual-slicing strategy. Specifically, CausalRepair employs context-aware static slicing on the test side to purify test semantics, and utilizes execution-trace-based dynamic slicing on the source side to capture precise runtime dependencies. This constructs a high-quality context causally relevant to the bug, which filters out irrelevant code and guides the iterative repair process. We evaluate CausalRepair on the widely used Defects4J (V1.2 and V2.0) and the latest Defects4J-Trans benchmarks. To ensure a fair comparison, we unify the backbone model as DeepSeek-V3 in all experiments. The results demonstrate that CausalRepair correctly fixes 313 bugs on Defects4J, significantly outperforming state-of-the-art approaches such as ReinFix and TSAPR, while reducing the average repair cost to $0.029 per bug, achieving a dual optimization of effectiveness and efficiency.
CCS Concepts: • Software and its engineering → Software testing and debugging.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on31
- CoCoNuT: combining context-aware neural translation models using ensemble for program repairThibaud Lutellier, Hung Viet Pham, Lawrence Pang, Yitong Li et al.ISSTA 2020 · 325 citations
- Automated Program Repair in the Era of Large Pre-trained Language ModelsChunqiu Steven Xia, Yuxiang Wei, Lingming ZhangICSE 2023 · 321 citations
- CURE: Code-Aware Neural Machine Translation for Automatic Program RepairNan Jiang, Thibaud Lutellier, Lin TanICSE 2021 · 267 citations
- Using an LLM to Help With Code UnderstandingDaye Nam, Andrew Macvean, Vincent J. Hellendoorn, Bogdan Vasilescu et al.ICSE 2024 · 264 citations
- Less training, more repairing please: revisiting automated program repair via zero-shot learningChunqiu Steven Xia, Lingming ZhangFSE 2022 · 223 citations
Related papers
- Automated Program Repair via Conversation: Fixing 162 out of 337 Bugs for $0.42 Each using ChatGPTChunqiu Steven Xia, Lingming ZhangISSTA 2024 · 105 citations
- Repair Ingredients Are All You Need: Improving Large Language Model-Based Program Repair via Repair Ingredients SearchJiayi Zhang, Kai Huang, Jian Zhang, Yang Liu et al.ICSE 2026
- ThinkRepair: Self-Directed Automated Program RepairXin Yin, Chao Ni, Shaohua Wang, Zhenhao Li et al.ISSTA 2024 · 37 citations
- A Deep Dive into Large Language Models for Automated Bug Localization and RepairSoneya Binta Hossain, Nan Jiang, Qiang Zhou, Xiaopeng Li et al.FSE 2024 · 60 citations
- RepairAgent: An Autonomous, LLM-Based Agent for Program RepairIslem Bouzenia, Premkumar T. Devanbu, Michael PradelICSE 2025 · 54 citations
