Towards Inferential Reproducibility of Machine Learning Research
Michael Hagmann, Philipp Meier, Stefan Riezler
Abstract
Reliability of machine learning evaluation -- the consistency of observed evaluation scores across replicated model training runs -- is affected by several sources of nondeterminism which can be regarded as measurement noise. Current tendencies to remove noise in order to enforce reproducibility of research results neglect inherent nondeterminism at the implementation level and disregard crucial interaction effects between algorithmic noise factors and data properties. This limits the scope of conclusions that can be drawn from such experiments. Instead of removing noise, we propose to incorporate several sources of variance, including their interaction with data properties, into an analysis of significance and reliability of machine learning evaluation, with the aim to draw inferences beyond particular instances of trained models. We show how to use linear mixed effects models (LMEMs) to analyze performance evaluation scores, and to conduct statistical inference with a generalized likelihood ratio test (GLRT). This allows us to incorporate arbitrary sources of noise like meta-parameter variations into statistical significance testing, and to assess performance differences conditional on data properties. Furthermore, a variance component analysis (VCA) enables the analysis of the contribution of noise sources to overall variance and the computation of a reliability coefficient by the ratio of substantial to total variance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 19501634-6ef8-4033-8030-fa835d9c2540Builds on7
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Extractive Summarization as Text MatchingMing Zhong, Pengfei Liu, Yiran Chen, Danqing Wang et al.ACL 2020 · 410 citations
- The MultiBERTs: BERT Reproductions for Robustness AnalysisThibault Sellam, Steve Yadlowsky, Ian Tenney, Jason Wei et al.ICLR 2022 · 106 citations
- Problems and Opportunities in Training Deep Learning Software Systems: An Analysis of VarianceHung Viet Pham, Shangshu Qian, Jiannan Wang, Thibaud Lutellier et al.ASE 2020 · 91 citations
- Reproducibility in Optimization: Theoretical Framework and LimitsKwangjun Ahn, Prateek Jain, Ziwei Ji, Satyen Kale et al.NeurIPS 2022 · 32 citations
Related papers
- Forest vs Tree: The (N, K) Trade-off in Reproducible ML EvaluationDeepak Pandita, Flip Korn, Chris Welty, Christopher M. HomanAAAI 2026 · 2 citations
- Quantified Reproducibility Assessment of NLP ResultsAnya Belz, Maja Popovic, Simon MilleACL 2022
- Signal and Noise: A Framework for Reducing Uncertainty in Language Model EvaluationDavid Heineman, Valentin Hofmann, Ian Magnusson, Yuling Gu et al.NeurIPS 2025 · 24 citations
- Rogue ScoresMax GruskyACL 2023 · 11 citations
- We Need to Talk About Reproducibility in NLP Model ComparisonYan Xue, Xuefei Cao, Xingli Yang, Yu Wang et al.EMNLP 2023 · 2 citations
