Community expectations for research artifacts and evaluation processes
Ben Hermann, Stefan Winter, Janet Siegmund
Abstract
Background. Artifact evaluation has been introduced into the software engineering and programming languages research community with a pilot at ESEC/FSE 2011 and has since then enjoyed a healthy adoption throughout the conference landscape. Objective. In this qualitative study, we examine the expectations of the community toward research artifacts and their evaluation processes. Method. We conducted a survey including all members of artifact evaluation committees of major conferences in the software engineering and programming language field since the first pilot and compared the answers to expectations set by calls for artifacts and reviewing guidelines. Results. While we find that some expectations exceed the ones expressed in calls and reviewing guidelines, there is no consensus on quality thresholds for artifacts in general. We observe very specific quality expectations for specific artifact types for review and later usage, but also a lack of their communication in calls. We also find problematic inconsistencies in the terminology used to express artifact evaluation’s most important purpose – replicability. Conclusion. We derive several actionable suggestions which can help to mature artifact evaluation in the inspected community and also to aid its introduction into other communities in computer science.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4a05eab6-05fb-46ff-b78a-54d8ccd7b4bdCited by top-tier papers4
- Boosting fuzzer efficiency: an information theoretic perspectiveMarcel Böhme, Valentin J. M. Manès, Sang Kil ChaFSE 2020 · 115 citations
- On the Reliability of Coverage-Based Fuzzer BenchmarkingMarcel Böhme, László Szekeres, Jonathan MetzmanICSE 2022 · 91 citations
- A retrospective study of one decade of artifact evaluationsStefan Winter, Christopher Steven Timperley, Ben Hermann, Jürgen Cito et al.FSE 2022 · 21 citations
- Not All Those Who Share Are Lost: Analyzing 25 Years of Cybersecurity Artifact Sharing Practices Through Automated DiscoveryDaan Vansteenhuyse, Arthur Bols, Lieven Desmet, Victor Le Pochat et al.USENIX Security 2026
Builds on1
Related papers
- How Transparent is Usable Privacy and Security Research? A Meta-Study on Current Research Transparency PracticesJan H. Klemmer, Juliane Schmüser, Fabian Fischer, Jacques Suray et al.USENIX Security 2025
- The State of Open Science in Software Engineering Research: A Case Study of ICSE ArtifactsAl Muttakin, Saikat Mondal, Chanchal K. RoyICSE 2026
- "Get in Researchers; We're Measuring Reproducibility": A Reproducibility Study of Machine Learning Papers in Tier 1 Security ConferencesDaniel Olszewski, Allison Lu, Carson Stillman, Kevin Warren et al.CCS 2023 · 19 citations
- Code replicability in computer graphicsNicolas Bonneel, David Coeurjolly, Julie Digne, Nicolas MelladoSIGGRAPH 2020 · 21 citations
- Views on Internal and External Validity in Empirical Software Engineering: 10 Years Later and BeyondAlina Mailach, Janet Siegmund, Sven Apel, Norbert SiegmundICSE 2026 · 1 citation
