Quantified Reproducibility Assessment of NLP Results
Anya Belz, Maja Popovic, Simon Mille
Abstract
This paper describes and tests a method for carrying out quantified reproducibility assessment (QRA) that is based on concepts and definitions from metrology. QRA produces a single score estimating the degree of reproducibility of a given system and evaluation measure, on the basis of the scores from, and differences between, different reproductions. We test QRA on 18 system and evaluation measure combinations (involving diverse NLP tasks and types of evaluation), for each of which we have the original results and one to seven reproduction results. The proposed QRA method produces degree-of-reproducibility scores that are comparable across multiple reproductions not only of the same, but of different original studies. We find that the proposed method facilitates insights into causes of variation between reproductions, and allows conclusions to be drawn about what changes to system and/or evaluation design might lead to improved reproducibility.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- KoLA: Carefully Benchmarking World Knowledge of Large Language ModelsJifan Yu, Xiaozhi Wang, Shangqing Tu, Shulin Cao et al.ICLR 2024 · 91 citations
- Reproducibility in Computational Linguistics: Is Source Code Enough?Mohammad Arvan, Luís Pina, Natalie PardeEMNLP 2022 · 12 citations
- Reproducibility Issues for BERT-based Evaluation MetricsYanran Chen, Jonas Belouadi, Steffen EgerEMNLP 2022 · 11 citations
- We Need to Talk About Reproducibility in NLP Model ComparisonYan Xue, Xuefei Cao, Xingli Yang, Yu Wang et al.EMNLP 2023 · 2 citations
- When Good and Reproducible Results are a Giant with Feet of Clay: The Importance of Software Quality in NLPSara Papi, Marco Gaido, Andrea Pilzer, Matteo NegriACL 2024
Related papers
- Rogue ScoresMax GruskyACL 2023 · 11 citations
- How to Measure the Reproducibility of System-oriented IR ExperimentsTimo Breuer, Nicola Ferro, Norbert Fuhr, Maria Maistro et al.SIGIR 2020 · 28 citations
- Towards Inferential Reproducibility of Machine Learning ResearchMichael Hagmann, Philipp Meier, Stefan RiezlerICLR 2023 · 1 citation
- Evaluating the Performance of Reinforcement Learning AlgorithmsScott M. Jordan, Yash Chandak, Daniel Cohen, Mengxue Zhang et al.ICML 2020 · 59 citations
- Scientific Credibility of Machine Translation Research: A Meta-Evaluation of 769 PapersBenjamin Marie, Atsushi Fujita, Raphael RubinoACL 2021
