When Good and Reproducible Results are a Giant with Feet of Clay: The Importance of Software Quality in NLP
Sara Papi, Marco Gaido, Andrea Pilzer, Matteo Negri
Abstract
Despite its crucial role in research experiments, code correctness is often presumed only on the basis of the perceived quality of results. This assumption comes with the risk of erroneous outcomes and potentially misleading findings. To address this issue, we posit that the current focus on reproducibility should go hand in hand with the emphasis on software quality. We present a case study in which we identify and fix three bugs in widely used implementations of the stateof-the-art Conformer architecture. Through experiments on speech recognition and translation in various languages, we demonstrate that the presence of bugs does not prevent the achievement of good and reproducible results, which however can lead to incorrect conclusions that potentially misguide future research. As a countermeasure, we propose a Code-quality Checklist and release pangoliNN, a library dedicated to testing neural models, with the goal of promoting coding best practices and improving research software quality within the NLP community.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8ff87552-4674-4e6b-8fe1-70bef985b4d3Cited by top-tier papers1
Ask how each one uses itBuilds on7
- Do Transformer Modifications Transfer Across Implementations and Applications?Sharan Narang, Hyung Won Chung, Yi Tay, Liam Fedus et al.EMNLP 2021 · 80 citations
- Beyond Accuracy: Behavioral Testing of NLP Models with CheckListMarco Túlio Ribeiro, Tongshuang Wu, Carlos Guestrin, Sameer SinghACL 2020 · 51 citations
- Revisiting End-to-End Speech-to-Text Translation From ScratchBiao Zhang, Barry Haddow, Rico SennrichICML 2022 · 46 citations
- Reproducibility in Computational Linguistics: Is Source Code Enough?Mohammad Arvan, Luís Pina, Natalie PardeEMNLP 2022 · 12 citations
- Stacked Acoustic-and-Textual Encoding: Integrating the Pre-trained Models into Speech Translation EncodersChen Xu, Bojie Hu, Yanyang Li, Yuhao Zhang et al.ACL 2021
Related papers
- Measuring the Influence of Incorrect Code on Test GenerationDong Huang, Jie M. Zhang, Mark Harman, Mingzhe Du et al.ICSE 2026
- Deep learning library testing via effective model generationZan Wang, Ming Yan, Junjie Chen, Shuang Liu et al.FSE 2020 · 165 citations
- When to Say What: Learning to Find Condition-Message InconsistenciesIslem Bouzenia, Michael PradelICSE 2023 · 5 citations
- Lost in Translation: A Study of Bugs Introduced by Large Language Models while Translating CodeRangeet Pan, Ali Reza Ibrahimzada, Rahul Krishna, Divya Sankar et al.ICSE 2024 · 96 citations
- Beyond Accuracy: Evaluating Self-Consistency of Code Large Language Models with IdentityChainMarcus J. Min, Yangruibo Ding, Luca Buratti, Saurabh Pujar et al.ICLR 2024 · 39 citations
