Uncontextualized significance considered dangerous
Nicola Ferro, Mark Sanderson
摘要
We examine the context of significance tests in offline retrieval experiments. Our Information Retrieval (IR) community is notable for its experimental rigour: the use of statistical significance is grows across our publications. However, we show that ignoring the context of a test risks Type I errors, leading to potential publication bias. We examine two contexts: multiple testing and the types of the retrieval systems being compared. Our results show that multiple testing corrections are critical for experimental work. In addition, we find that past research on the reliability of test collections maybe flawed owing to the type of systems examined. The latter result has not been shown before. Together our results suggest substantial numbers of Type I errors in offline IR experiments. We detail a methodology to alleviate the errors.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- Agreement and Disagreement between True and False-Positive Metrics in Recommender Systems EvaluationElisa Mena-Maldonado, Rocío Cañamares, Pablo Castells, Yongli Ren 等SIGIR 2020 · 被引用 16 次
- How to Measure the Reproducibility of System-oriented IR ExperimentsTimo Breuer, Nicola Ferro, Norbert Fuhr, Maria Maistro 等SIGIR 2020 · 被引用 28 次
- Neural Retrievers are Biased Towards LLM-Generated ContentSunhao Dai, Yuqi Zhou, Liang Pang, Weihao Liu 等KDD 2024 · 被引用 26 次
- Bayesian Inferential Risk Evaluation On Multiple IR SystemsRodger Benham, Ben Carterette, J. Shane Culpepper, Alistair MoffatSIGIR 2020 · 被引用 5 次
- Invisible Relevance Bias: Text-Image Retrieval Models Prefer AI-Generated ImagesShicheng Xu, Danyang Hou, Liang Pang, Jingcheng Deng 等SIGIR 2024 · 被引用 18 次
