Reliable Confidence Intervals for Information Retrieval Evaluation Using Generative A.I
Harrie Oosterhuis, Rolf Jagerman, Zhen Qin, Xuanhui Wang, Michael Bendersky
Abstract
The traditional evaluation of information retrieval (IR) systems is generally very costly as it requires manual relevance annotation from human experts. Recent advancements in generative artificial intelligence -specifically large language models (LLMs)-can generate relevance annotations at an enormous scale with relatively small computational costs. Potentially, this could alleviate the costs traditionally associated with IR evaluation and make it applicable to numerous low-resource applications. However, generated relevance annotations are not immune to (systematic) errors, and as a result, directly using them for evaluation produces unreliable results.
In this work, we propose two methods based on predictionpowered inference and conformal risk control that utilize computergenerated relevance annotations to place reliable confidence intervals (CIs) around IR evaluation metrics. Our proposed methods require a small number of reliable annotations from which the methods can statistically analyze the errors in the generated annotations. Using this information, we can place CIs around evaluation metrics with strong theoretical guarantees. Unlike existing approaches, our conformal risk control method is specifically designed for ranking metrics and can vary its CIs per query and document. Our experimental results show that our CIs accurately capture both the variance and bias in evaluation based on LLM annotations, better than the typical empirical bootstrapping estimates. We hope our contributions bring reliable evaluation to the many IR applications where this was traditionally infeasible.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e85fa4de-f45a-4527-b6eb-08b77f2c8007Cited by top-tier papers4
- Don’t Pass@k: A Bayesian Framework for Large Language Model EvaluationMohsen Hariri, Amirhossein Samandar, Michael Hinczewski, Vipin ChaudharyICLR 2026 · 18 citations
- General Synthetic-Powered InferenceMeshi Bashari, Yonghoon Lee, Roy Lotan, Edgar Dobriban et al.ICML 2026 · 5 citations
- RSA-CP: Efficient Conformal Prediction in Small-Sample Regimes via Random Score AlignmentPankaj Bhagwat, Zhixian Yang, yihao wang, Bei Jiang et al.ICML 2026
- AGRAG: Advanced Graph-Based Retrieval-Augmented Generation for LLMsYubo Wang, Haoyang Li, Fei Teng, Lei ChenICDE 2026
Builds on6
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- The Flan Collection: Designing Data and Methods for Effective Instruction TuningShayne Longpre, Le Hou, Tu Vu, Albert Webson et al.ICML 2023 · 908 citations
- Conformal Risk ControlAnastasios Nikolas Angelopoulos, Stephen Bates, Adam Fisch, Lihua Lei et al.ICLR 2024 · 242 citations
- Large Language Models can Accurately Predict Searcher PreferencesPaul Thomas, Seth Spielman, Nick Craswell, Bhaskar MitraSIGIR 2024 · 153 citations
Related papers
- Conf-Gen: Conformal Uncertainty Quantification for Generative ModelsGabriel Loaiza-Ganem, Kevin Zhang, Wei Cui, Marc Law et al.ICML 2026 · 1 citation
- Analyzing Uncertainty of LLM-as-a-Judge: Interval Evaluations with Conformal PredictionHuanxin Sheng, Xinyi Liu, Hangfeng He, Jieyu Zhao et al.EMNLP 2025 · 1 citation
- Conformal Reliability: A New Evaluation Metric for Conditional GenerationYachen Gao, Xinwei Sun, Yikai Wang, Ye Shi et al.ICML 2026
- C-RAG: Certified Generation Risks for Retrieval-Augmented Language ModelsMintong Kang, Nezihe Merve Gürel, Ning Yu, Dawn Song et al.ICML 2024 · 33 citations
- Neural Retrievers are Biased Towards LLM-Generated ContentSunhao Dai, Yuqi Zhou, Liang Pang, Weihao Liu et al.KDD 2024 · 26 citations
