Reproducibility in Computational Linguistics: Is Source Code Enough?
Mohammad Arvan, Luís Pina, Natalie Parde
Abstract
The availability of source code has been put forward as one of the most critical factors for improving the reproducibility of scientific research. This work studies trends in source code availability at major computational linguistics conferences, namely, ACL, EMNLP, LREC, NAACL, and COLING. We observe positive trends, especially in conferences that actively promote reproducibility. We follow this by conducting a reproducibility study of eight papers published in EMNLP 2021, finding that source code releases leave much to be desired. Moving forward, we suggest all conferences require self-contained artifacts and provide a venue to evaluate such artifacts at the time of publication. Authors can include small-scale experiments and explicit scripts to generate each result to improve the reproducibility of their work.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 739661cc-c78c-4d8b-a377-a7a66ec9bfb2Cited by top-tier papers6
- Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language ModelsWenxuan Wang, Zizhan Ma, Guo Yu, Yiu-Fai Cheung et al.ACL 2026 · 9 citations
- The Unreasonable Effectiveness of Open Science in AI: A Replication StudyOdd Erik Gundersen, Odd Cappelen, Martin Mølnå, Nicklas Grimstad NilsenAAAI 2025 · 8 citations
- We Need to Talk About Reproducibility in NLP Model ComparisonYan Xue, Xuefei Cao, Xingli Yang, Yu Wang et al.EMNLP 2023 · 2 citations
- Forest vs Tree: The (N, K) Trade-off in Reproducible ML EvaluationDeepak Pandita, Flip Korn, Chris Welty, Christopher M. HomanAAAI 2026 · 2 citations
- GSAP-ERE: Fine-Grained Scholarly Entity and Relation Extraction Focused on Machine LearningWolfgang Otto, Lu Gan, Sharmila Upadhyaya, Saurav Karmakar et al.AAAI 2026
Builds on8
- Weakly-supervised Text Classification Based on Keyword GraphLu Zhang, Jiandong Ding, Yi Xu, Yingyao Liu et al.EMNLP 2021 · 46 citations
- StreamHover: Livestream Transcript Summarization and AnnotationSangwoo Cho, Franck Dernoncourt, Tim Ganter, Trung Bui et al.EMNLP 2021 · 18 citations
- Measuring Association Between Labels and Free-Text RationalesSarah Wiegreffe, Ana Marasovic, Noah A. SmithEMNLP 2021 · 12 citations
- Automatically Exposing Problems with Neural Dialog ModelsDian Yu, Kenji SagaeEMNLP 2021 · 5 citations
- A Massively Multilingual Analysis of Cross-linguality in Shared Embedding SpaceAlexander Jones, William Yang Wang, Kyle MahowaldEMNLP 2021 · 4 citations
Related papers
- NLP Reproducibility For All: Understanding Experiences of BeginnersShane Storks, Keunwoo Peter Yu, Ziqiao Ma, Joyce ChaiACL 2023
- Code replicability in computer graphicsNicolas Bonneel, David Coeurjolly, Julie Digne, Nicolas MelladoSIGGRAPH 2020 · 21 citations
- "Get in Researchers; We're Measuring Reproducibility": A Reproducibility Study of Machine Learning Papers in Tier 1 Security ConferencesDaniel Olszewski, Allison Lu, Carson Stillman, Kevin Warren et al.CCS 2023 · 19 citations
- Reflections on the Reproducibility of Commercial LLM Performance in Empirical Software Engineering StudiesFlorian Angermeir, Maximilian Amougou, Mark Kreitz, Andreas Bauer et al.ICSE 2026 · 1 citation
- The State of Open Science in Software Engineering Research: A Case Study of ICSE ArtifactsAl Muttakin, Saikat Mondal, Chanchal K. RoyICSE 2026
