Quantified Reproducibility Assessment of NLP Results
Anya Belz, Maja Popovic, Simon Mille
摘要
This paper describes and tests a method for carrying out quantified reproducibility assessment (QRA) that is based on concepts and definitions from metrology. QRA produces a single score estimating the degree of reproducibility of a given system and evaluation measure, on the basis of the scores from, and differences between, different reproductions. We test QRA on 18 system and evaluation measure combinations (involving diverse NLP tasks and types of evaluation), for each of which we have the original results and one to seven reproduction results. The proposed QRA method produces degree-of-reproducibility scores that are comparable across multiple reproductions not only of the same, but of different original studies. We find that the proposed method facilitates insights into causes of variation between reproductions, and allows conclusions to be drawn about what changes to system and/or evaluation design might lead to improved reproducibility.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- KoLA: Carefully Benchmarking World Knowledge of Large Language ModelsJifan Yu, Xiaozhi Wang, Shangqing Tu, Shulin Cao 等ICLR 2024 · 被引用 91 次
- Reproducibility in Computational Linguistics: Is Source Code Enough?Mohammad Arvan, Luís Pina, Natalie PardeEMNLP 2022 · 被引用 12 次
- Reproducibility Issues for BERT-based Evaluation MetricsYanran Chen, Jonas Belouadi, Steffen EgerEMNLP 2022 · 被引用 11 次
- We Need to Talk About Reproducibility in NLP Model ComparisonYan Xue, Xuefei Cao, Xingli Yang, Yu Wang 等EMNLP 2023 · 被引用 2 次
- When Good and Reproducible Results are a Giant with Feet of Clay: The Importance of Software Quality in NLPSara Papi, Marco Gaido, Andrea Pilzer, Matteo NegriACL 2024
相关 Paper
- Rogue ScoresMax GruskyACL 2023 · 被引用 11 次
- How to Measure the Reproducibility of System-oriented IR ExperimentsTimo Breuer, Nicola Ferro, Norbert Fuhr, Maria Maistro 等SIGIR 2020 · 被引用 28 次
- Towards Inferential Reproducibility of Machine Learning ResearchMichael Hagmann, Philipp Meier, Stefan RiezlerICLR 2023 · 被引用 1 次
- Evaluating the Performance of Reinforcement Learning AlgorithmsScott M. Jordan, Yash Chandak, Daniel Cohen, Mengxue Zhang 等ICML 2020 · 被引用 59 次
- Scientific Credibility of Machine Translation Research: A Meta-Evaluation of 769 PapersBenjamin Marie, Atsushi Fujita, Raphael RubinoACL 2021
