How to Measure the Reproducibility of System-oriented IR Experiments
Timo Breuer, Nicola Ferro, Norbert Fuhr, Maria Maistro, Tetsuya Sakai, Philipp Schaer, Ian Soboroff
摘要
Replicability and reproducibility of experimental results are primary concerns in all the areas of science and IR is not an exception. Besides the problem of moving the field towards more reproducible experimental practices and protocols, we also face a severe methodological issue: we do not have any means to assess when reproduced is reproduced. Moreover, we lack any reproducibility-oriented dataset, which would allow us to develop such methods.
To address these issues, we compare several measures to objectively quantify to what extent we have replicated or reproduced a system-oriented IR experiment. These measures operate at different levels of granularity, from the fine-grained comparison of ranked lists, to the more general comparison of the obtained effects and significant differences. Moreover, we also develop a reproducibilityoriented dataset, which allows us to validate our measures and which can also be used to develop future measures.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
相关 Paper
- Quantified Reproducibility Assessment of NLP ResultsAnya Belz, Maja Popovic, Simon MilleACL 2022
- Reproduce, Replicate, Reevaluate. The Long but Safe Way to Extend Machine Learning MethodsLuisa Werner, Nabil Layaïda, Pierre Genevès, Jérôme Euzenat 等AAAI 2024 · 被引用 2 次
- How Machine Learning Is Solving the Binary Function Similarity ProblemAndrea Marcelli, Mariano Graziano, Xabier Ugarte-Pedrero, Yanick Fratantonio 等USENIX Security 2022
- Coordinating Chaos: A Structured Review of Linguistic Coordination MethodologiesBenjamin Roger Litterer, David Jurgens, Dallas CardACL 2025
- A Fair Comparison of Graph Neural Networks for Graph ClassificationFederico Errica, Marco Podda, Davide Bacciu, Alessio MicheliICLR 2020 · 被引用 508 次
