Towards Inferential Reproducibility of Machine Learning Research
Michael Hagmann, Philipp Meier, Stefan Riezler
摘要
Reliability of machine learning evaluation -- the consistency of observed evaluation scores across replicated model training runs -- is affected by several sources of nondeterminism which can be regarded as measurement noise. Current tendencies to remove noise in order to enforce reproducibility of research results neglect inherent nondeterminism at the implementation level and disregard crucial interaction effects between algorithmic noise factors and data properties. This limits the scope of conclusions that can be drawn from such experiments. Instead of removing noise, we propose to incorporate several sources of variance, including their interaction with data properties, into an analysis of significance and reliability of machine learning evaluation, with the aim to draw inferences beyond particular instances of trained models. We show how to use linear mixed effects models (LMEMs) to analyze performance evaluation scores, and to conduct statistical inference with a generalized likelihood ratio test (GLRT). This allows us to incorporate arbitrary sources of noise like meta-parameter variations into statistical significance testing, and to assess performance differences conditional on data properties. Furthermore, a variance component analysis (VCA) enables the analysis of the contribution of noise sources to overall variance and the computation of a reliability coefficient by the ratio of substantial to total variance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Extractive Summarization as Text MatchingMing Zhong, Pengfei Liu, Yiran Chen, Danqing Wang 等ACL 2020 · 被引用 410 次
- The MultiBERTs: BERT Reproductions for Robustness AnalysisThibault Sellam, Steve Yadlowsky, Ian Tenney, Jason Wei 等ICLR 2022 · 被引用 106 次
- Problems and Opportunities in Training Deep Learning Software Systems: An Analysis of VarianceHung Viet Pham, Shangshu Qian, Jiannan Wang, Thibaud Lutellier 等ASE 2020 · 被引用 91 次
- Reproducibility in Optimization: Theoretical Framework and LimitsKwangjun Ahn, Prateek Jain, Ziwei Ji, Satyen Kale 等NeurIPS 2022 · 被引用 32 次
相关 Paper
- Forest vs Tree: The (N, K) Trade-off in Reproducible ML EvaluationDeepak Pandita, Flip Korn, Chris Welty, Christopher M. HomanAAAI 2026 · 被引用 2 次
- Quantified Reproducibility Assessment of NLP ResultsAnya Belz, Maja Popovic, Simon MilleACL 2022
- Signal and Noise: A Framework for Reducing Uncertainty in Language Model EvaluationDavid Heineman, Valentin Hofmann, Ian Magnusson, Yuling Gu 等NeurIPS 2025 · 被引用 24 次
- Rogue ScoresMax GruskyACL 2023 · 被引用 11 次
- We Need to Talk About Reproducibility in NLP Model ComparisonYan Xue, Xuefei Cao, Xingli Yang, Yu Wang 等EMNLP 2023 · 被引用 2 次
