A retrospective study of one decade of artifact evaluations
Stefan Winter, Christopher Steven Timperley, Ben Hermann, Jürgen Cito, Jonathan Bell, Michael Hilton, Dirk Beyer
Abstract
Most software engineering research involves the development of a prototype, a proof of concept, or a measurement apparatus. Together with the data collected in the research process, they are collectively referred to as research artifacts and are subject to artifact evaluation (AE) at scientific conferences. Since its initiation in the SE community at ESEC/FSE 2011, both the goals and the process of AE have evolved and today expectations towards AE are strongly linked with reproducible research results and reusable tools that other researchers can build their work on. However, to date little evidence has been provided that artifacts which have passed AE actually live up to these high expectations, i.e., to which degree AE processes contribute to AE's goals and whether the overhead they impose is justified.
We aim to fill this gap by providing an in-depth analysis of research artifacts from a decade of software engineering (SE) and programming languages (PL) conferences, based on which we reflect on the goals and mechanisms of AE in our community. In summary, our analyses (1) suggest that articles with artifacts do not generally have better visibility in the community, (2) provide evidence how evaluated and not evaluated artifacts differ with respect to different quality criteria, and (3) highlight opportunities for further improving AE processes.
• General and reference → Empirical studies; • Software and its engineering → Software post-development issues; • Information systems → Digital libraries and archives.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 914ec7de-6b87-47c0-a6e1-44cd1719c24dCited by top-tier papers2
- "Get in Researchers; We're Measuring Reproducibility": A Reproducibility Study of Machine Learning Papers in Tier 1 Security ConferencesDaniel Olszewski, Allison Lu, Carson Stillman, Kevin Warren et al.CCS 2023 · 19 citations
- Not All Those Who Share Are Lost: Analyzing 25 Years of Cybersecurity Artifact Sharing Practices Through Automated DiscoveryDaan Vansteenhuyse, Arthur Bols, Lieven Desmet, Victor Le Pochat et al.USENIX Security 2026
Builds on3
- Transparency of CHI Research Artifacts: Results of a Self-Reported SurveyChat Wacharamanotham, Lukas Eisenring, Steve Haroz, Florian EchtlerCHI 2020 · 120 citations
- Community expectations for research artifacts and evaluation processesBen Hermann, Stefan Winter, Janet SiegmundFSE 2020 · 40 citations
- Code replicability in computer graphicsNicolas Bonneel, David Coeurjolly, Julie Digne, Nicolas MelladoSIGGRAPH 2020 · 21 citations
Related papers
- The State of Open Science in Software Engineering Research: A Case Study of ICSE ArtifactsAl Muttakin, Saikat Mondal, Chanchal K. RoyICSE 2026
- Reproducibility in Computational Linguistics: Is Source Code Enough?Mohammad Arvan, Luís Pina, Natalie PardeEMNLP 2022 · 12 citations
- Sharing Software-Evolution Datasets: Practices, Challenges, and RecommendationsDavid Broneske, Sebastian Kittan, Jacob KrügerFSE 2024 · 4 citations
- Data to Infinity and Beyond: Examining Data Sharing and Reuse Practices in the Computer Security CommunityAnna Crowder, Allison Lu, Kevin Childs, Carson Stillman et al.S&P 2025
- How Transparent is Usable Privacy and Security Research? A Meta-Study on Current Research Transparency PracticesJan H. Klemmer, Juliane Schmüser, Fabian Fischer, Jacques Suray et al.USENIX Security 2025
