Names Are All You Need: Effective and Safe Regression Test Selection for Python
You Wang, Michael Pradel, Zhongxin Liu
摘要
Regression test selection (RTS) reduces the cost of regression testing by executing only those tests affected by a code change. Despite extensive study of RTS in statically typed languages such as Java, achieving effective and safe RTS in Python is challenging. Python’s dynamic typing makes precise call-graph construction difficult, which can cause call-graph-based RTS to miss affected tests, and hence, compromise safety. Python’s eager importing mechanism, in contrast, renders file-level dependency analysis overly conservative. This paper presents NameRTS, the first Python RTS approach based on fine-grained dependency analysis. NameRTS models a Python program as a bipartite graph of code element nodes (e.g., classes, functions, global variables) and name nodes (i.e., identifiers used to reference code elements), with edges capturing definitions and references. RTS is formulated as a reachability problem on this graph: a test is selected if any modified code element is reachable from the names used in that test. This design avoids call-graph construction, enabling a conservative analysis amenable to safety. To control dependency cascades introduced by coarse name matching, NameRTS applies two pruning strategies that leverage prior test executions and context information to refine name matching. To evaluate NameRTS, we construct the first Python RTS dataset with a ground truth indicating which test files are affected by each commit. It includes 500 commits drawn from 10 real-world Python projects. We compare NameRTS with the best-performing baseline, BabelRTS, an RTS technique based on coarse file-level dependencies. On this benchmark, NameRTS skips 69.90% of test files on average, outperforming BabelRTS by 146.5%. It also reduces end-to-end testing time by 45.59%, yielding a 107.7% improvement over BabelRTS. In terms of safety, NameRTS selects all affected tests for 99.6% of commits, with only rare misses in exceptional cases. In contrast, BabelRTS is safe for 76.6% of commits. These results demonstrate the effectiveness of NameRTS, paving the way for more efficient regression testing in Python.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
- Learning-to-rank vs ranking-to-learn: strategies for regression testing in continuous integrationAntonia Bertolino, Antonio Guerriero, Breno Miranda, Roberto Pietrantuono 等ICSE 2020 · 被引用 81 次
- DynaPyt: a dynamic analysis framework for PythonAryaz Eghbali, Michael PradelFSE 2022 · 被引用 30 次
- Continuous test suite failure predictionCong Pan, Michael PradelISSTA 2021 · 被引用 22 次
- More Precise Regression Test Selection via Reasoning about Semantics-Modifying ChangesYu Liu, Jiyang Zhang, Pengyu Nie, Milos Gligoric 等ISSTA 2023 · 被引用 19 次
相关 Paper
- GameRTS: A Regression Testing Framework for Video GamesJiongchi Yu, Yuechen Wu, Xiaofei Xie, Wei Le 等ICSE 2023 · 被引用 5 次
- Hybrid Regression Test Selection by Integrating File and Method DependencesGuofeng Zhang, Luyao Liu, Zhenbang Chen, Ji WangASE 2024 · 被引用 2 次
- Reflective Unit Test Generation for Precise Type Error Detection with Large Language ModelsChen Yang, Ziqi Wang, Yanjie Jiang, Lin Yang 等ASE 2025 · 被引用 1 次
- The evolution of type annotations in python: an empirical studyLuca Di Grazia, Michael PradelFSE 2022 · 被引用 29 次
- Test Selection for Unified Regression TestingShuai Wang, Xinyu Lian, Darko Marinov, Tianyin XuICSE 2023 · 被引用 9 次
