ChangeGuard: Validating Code Changes via Pairwise Learning-Guided Execution
Lars Gröninger, Beatriz Souza, Michael Pradel
Abstract
Code changes are an integral part of the software development process. Many code changes are meant to improve the code without changing its functional behavior, e.g., refactorings and performance improvements. Unfortunately, validating whether a code change preserves the behavior is non-trivial, particularly when the code change is performed deep inside a complex project. This paper presents ChangeGuard, an approach that uses learning-guided execution to compare the runtime behavior of a modified function. The approach is enabled by the novel concept of pairwise learning-guided execution and by a set of techniques that improve the robustness and coverage of the state-of-the-art learning-guided execution technique. Our evaluation applies ChangeGuard to a dataset of 224 manually annotated code changes from popular Python open-source projects and to three datasets of code changes obtained by applying automated code transformations. Our results show that the approach identifies semantics-changing code changes with a precision of 77.1% and a recall of 69.5%, and that it detects unexpected behavioral changes introduced by automatic code refactoring tools. In contrast, the existing regression tests of the analyzed projects miss the vast majority of semantics-changing code changes, with a recall of only 7.6%. We envision our approach being useful for detecting unintended behavioral changes early in the development process and for improving the quality of automated code transformations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Testora: Using Natural Language Intent to Detect Behavioral RegressionsMichael PradelICSE 2026 · 1 citation
- Closing the Loop: Universal Repository Representation with RPG-EncoderJane Luo, Chengyu Yin, Xin Zhang, Qingtao Li et al.ICML 2026
- RippleGUItester: Change-Aware Exploratory TestingYanqi Su, Michael Pradel, Chunyang ChenISSTA 2026
- Treefix: Enabling Execution with a Tree of PrefixesBeatriz Souza, Michael PradelICSE 2025
Builds on9
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 1,224 citations
- CodaMosa: Escaping Coverage Plateaus in Test Generation with Pre-trained Large Language ModelsCaroline Lemieux, Jeevana Priya Inala, Shuvendu K. Lahiri, Siddhartha SenICSE 2023 · 221 citations
- CC2Vec: distributed representations of code changesThong Hoang, Hong Jin Kang, David Lo, Julia LawallICSE 2020 · 169 citations
- Code-Aware Prompting: A Study of Coverage-Guided Test Generation in Regression Setting using LLMGabriel Ryan, Siddhartha Jain, Mingyue Shang, Shiqi Wang et al.FSE 2024 · 68 citations
- Making Python code idiomatic by automatic refactoring non-idiomatic Python code with pythonic idiomsZejun Zhang, Zhenchang Xing, Xin Xia, Xiwei Xu et al.FSE 2022 · 35 citations
Related papers
- FlakeFlagger: Predicting Flakiness Without Rerunning TestsAbdulrahman Alshammari, Christopher Morris, Michael Hilton, Jonathan BellICSE 2021 · 63 citations
- Speculative Automated Refactoring of Imperative Deep Learning Programs to Graph ExecutionRaffi Khatchadourian, Tatiana Castro Vélez, Mehdi Bagherzadeh, Nan Jia et al.ASE 2025 · 1 citation
- PYEVOLVE: Automating Frequent Code Changes in Python ML SystemsMalinda Dilhara, Danny Dig, Ameya KetkarICSE 2023 · 46 citations
- Efficient Detection of Test Interference in C ProjectsFlorian Eder, Stefan WinterASE 2024 · 2 citations
- Reflective Unit Test Generation for Precise Type Error Detection with Large Language ModelsChen Yang, Ziqi Wang, Yanjie Jiang, Lin Yang et al.ASE 2025 · 1 citation
