Evidence Profiles for Validity Threats in Program Comprehension Experiments
Marvin Muñoz Barón, Marvin Wyrich, Daniel Graziotin, Stefan Wagner
Abstract
Searching for clues, gathering evidence, and reviewing case files are all techniques used by criminal investigators to draw sound conclusions and avoid wrongful convictions. Medicine, too, has a long tradition of evidence-based practice, in which administering a treatment without evidence of its efficacy is considered malpractice. Similarly, in software engineering (SE) research, we can develop sound methodologies and mitigate threats to validity by basing study design decisions on evidence. Echoing a recent call for the empirical evaluation of design decisions in program comprehension experiments, we conducted a 2-phases study consisting of systematic literature searches, snowballing, and thematic synthesis. We found out (1) which validity threat categories are most often discussed in primary studies of code comprehension, and we collected evidence to build (2) the evidence profiles for the three most commonly reported threats to validity. We discovered that few mentions of validity threats in primary studies (31 of 409) included a reference to supporting evidence. For the three most commonly mentioned threats, namely the influence of programming experience, program length, and the selected comprehension measures, almost all cited studies (17 of 18) did not meet our criteria for evidence. We show that for many threats to validity that are currently assumed to be influential across all studies, their actual impact may depend on the design and context of each specific study. Researchers should discuss threats to validity within the context of their particular study and support their discussions with evidence. The present paper can be one resource for evidence, and we call for more meta-studies of this type to be conducted, which will then inform design decisions in primary studies. Further, although we have applied our methodology in the context of program comprehension, our approach can also be used in other SE research areas to enable evidence-based experiment design decisions and meaningful discussions of threats to validity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itRelated papers
- "Threat modeling is very formal, it's very technical, and also very hard to do correctly": Investigating Threat Modeling Practices in Open-Source Software ProjectsHarjot Kaur, Carson Powers, Ronald E. Thompson III, Sascha Fahl et al.USENIX Security 2025
- An evidence-based inquiry into the use of grey literature in software engineeringHe Zhang, Xin Zhou, Xin Huang, Huang Huang et al.ICSE 2020 · 29 citations
- On the Relationship between Code Verifiability and UnderstandabilityKobi Feldman, Martin Kellogg, Oscar ChaparroFSE 2023 · 2 citations
- MetaExplorer : Facilitating Reasoning with Epistemic Uncertainty in Meta-analysisAlex Kale, Sarah Lee, Terrance Goan, Elizabeth Tipton et al.CHI 2023 · 8 citations
- Program Comprehension and Code Complexity Metrics: An fMRI StudyNorman Peitek, Sven Apel, Chris Parnin, André Brechmann et al.ICSE 2021 · 59 citations
