We Need to Talk About Reproducibility in NLP Model Comparison
Yan Xue, Xuefei Cao, Xingli Yang, Yu Wang, Ruibo Wang, Jihong Li
Abstract
NLPers frequently face reproducibility crisis in a comparison of various models of a real-world NLP task. Many studies have empirically showed that the standard splits tend to produce low reproducible and unreliable conclusions, and they attempted to improve the splits by using more random repetitions. However, the improvement on the reproducibility in a comparison of NLP models is limited attributed to a lack of investigation on the relationship between the reproducibility and the estimator induced by a splitting strategy. In this paper, we formulate the reproducibility in a model comparison into a probabilistic function with regard to a conclusion. Furthermore, we theoretically illustrate that the reproducibility is qualitatively dominated by the signal-to-noise ratio (SNR) of a model performance estimator obtained on a corpus splitting strategy. Specifically, a higher value of the SNR of an estimator probably indicates a better reproducibility. On the basis of the theoretical motivations, we develop a novel mixture estimator of the performance of an NLP model with a regularized corpus splitting strategy based on a blocked 3× 2 cross-validation. We conduct numerical experiments on multiple NLP tasks to show that the proposed estimator achieves a high SNR, and it substantially increases the reproducibility. Therefore, we recommend the NLP practitioners to use the proposed method to compare NLP models instead of the methods based on the widely-used standard splits and the random splits with multiple repetitions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b3cbf2f9-fb09-402a-bbf3-6ce8826024e8Cited by top-tier papers1
Ask how each one uses itBuilds on2
Related papers
- Signal and Noise: A Framework for Reducing Uncertainty in Language Model EvaluationDavid Heineman, Valentin Hofmann, Ian Magnusson, Yuling Gu et al.NeurIPS 2025 · 24 citations
- Towards Inferential Reproducibility of Machine Learning ResearchMichael Hagmann, Philipp Meier, Stefan RiezlerICLR 2023 · 1 citation
- NLP Reproducibility For All: Understanding Experiences of BeginnersShane Storks, Keunwoo Peter Yu, Ziqiao Ma, Joyce ChaiACL 2023
- When Good and Reproducible Results are a Giant with Feet of Clay: The Importance of Software Quality in NLPSara Papi, Marco Gaido, Andrea Pilzer, Matteo NegriACL 2024
- Reproducibility Issues for BERT-based Evaluation MetricsYanran Chen, Jonas Belouadi, Steffen EgerEMNLP 2022 · 11 citations
