Revisiting the Relationship Between Fault Detection, Test Adequacy Criteria, and Test Set Size
Yiqun T. Chen, Rahul Gopinath, Anita Tadakamalla, Michael D. Ernst, Reid Holmes, Gordon Fraser, Paul Ammann, René Just
Abstract
The research community has long recognized a complex interrelationship between fault detection, test adequacy criteria, and test set size. However, there is substantial confusion about whether and how to experimentally control for test set size when assessing how well an adequacy criterion is correlated with fault detection and when comparing test adequacy criteria. Resolving the confusion, this paper makes the following contributions: (1) A review of contradictory analyses of the relationships between fault detection, test adequacy criteria, and test set size. Specifically, this paper addresses the supposed contradiction of prior work and explains why test set size is neither a confounding variable, as previously suggested, nor an independent variable that should be experimentally manipulated. (2) An explication and discussion of the experimental designs of prior work, together with a discussion of conceptual and statistical problems, as well as specific guidelines for future work. (3) A methodology for comparing test adequacy criteria on an equal basis, which accounts for test set size without directly manipulating it through unrealistic stratification. (4) An empirical evaluation that compares the effectiveness of coverage-based testing, mutation-based testing, and random testing. Additionally, this paper proposes probabilistic coupling, a methodology for assessing the representativeness of a set of test goals for a given fault and for approximating the fault-detection probability of adequate test sets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 97fe7464-6dc7-4d0e-b9bc-b65a438416d7Cited by top-tier papers11
- On the Reliability of Coverage-Based Fuzzer BenchmarkingMarcel Böhme, László Szekeres, Jonathan MetzmanICSE 2022 · 91 citations
- Prioritizing Mutants to Guide Mutation TestingSamuel J. Kaufman, Ryan Featherman, Justin Alvin, Bob Kurtz et al.ICSE 2022 · 36 citations
- Guiding Greybox Fuzzing with Mutation TestingVasudev Vikram, Isabella Laybourn, Ao Li, Nicole Nair et al.ISSTA 2023 · 22 citations
- Growing A Test Corpus with Bonsai FuzzingVasudev Vikram, Rohan Padhye, Koushik SenICSE 2021 · 15 citations
- Reachable Coverage: Estimating Saturation in FuzzingDanushka Liyanage, Marcel Böhme, Chakkrit Tantithamthavorn, Stephan LippICSE 2023 · 14 citations
Related papers
- Does mutation testing improve testing practices?Goran Petrovic, Marko Ivankovic, Gordon Fraser, René JustICSE 2021 · 6 citations
- Measuring and Mitigating Gaps in Structural TestingSoneya Binta Hossain, Matthew B. Dwyer, Sebastian G. Elbaum, Anh Nguyen-TuongICSE 2023 · 7 citations
- Systematic Assessment of Fuzzers using Mutation AnalysisPhilipp Görz, Björn Mathis, Keno Hassler, Emre Güler et al.USENIX Security 2023
- Navigating Mobile Testing Evaluation: A Comprehensive Statistical Analysis of Android GUI Testing MetricsYuanhong Lan, Yifei Lu, Minxue Pan, Xuandong LiASE 2024 · 4 citations
- To Kill a Mutant: An Empirical Study of Mutation Testing KillsHang Du, Vijay Krishna Palepu, James A. JonesISSTA 2023 · 5 citations
