Uncertainty Quantification for Retrieval-Augmented Reasoning
Heydar Soudani, Hamed Zamani, Faegheh Hasibi
Abstract
Retrieval-augmented reasoning (RAR) is a recent evolution of retrieval-augmented generation (RAG) that employs multiple reasoning steps for retrieval and generation. While effective for some complex queries, RAR remains vulnerable to errors and misleading outputs. Uncertainty quantification (UQ) offers methods to estimate the confidence of systems' outputs. These methods, however, often handle simple queries with no retrieval or single-step retrieval, without properly handling RAR setup. Accurate estimation of UQ for RAR requires accounting for all sources of uncertainty, including those arising from retrieval and generation. In this paper, we account for these sources and introduce Retrieval-Augmented Reasoning Consistency (R 2 C), a novel UQ method for RAR. The core idea of R 2 C is to perturb the multi-step reasoning process by applying various actions to reasoning steps. These perturbations alter the retriever's input, which shifts its output and consequently modifies the generator's input at the next step. Through this iterative feedback loop, the retriever and generator continuously reshape each other's inputs, enabling us to capture uncertainty arising from both components. Experiments on five popular RAR systems across diverse QA datasets show that R 2 C improves AU-ROC by over 5% on average compared to the state-of-the-art UQ baselines. Extrinsic evaluations using R 2 C as an external signal further confirm its effectiveness for two downstream tasks: in the Abstention task, it achieves 5% gains in both F1Abstain and Ac-cAbstain; in Model Selection, it improves exact match by 7% over single models and 3% over selection methods. Code is available on https://github.com/HeydarSoudani/R2C.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f6cf3586-e8c7-4597-9897-2b7ae30fe339Cited by top-tier papers1
Ask how each one uses itBuilds on34
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil et al.ICLR 2024 · 1,798 citations
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMsMiao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li et al.ICLR 2024 · 867 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
- Uncertainty Estimation in Autoregressive Structured PredictionAndrey Malinin, Mark J. F. GalesICLR 2021 · 439 citations
- Active Retrieval Augmented GenerationZhengbao Jiang, Frank F. Xu, Luyu Gao, Zhiqing Sun et al.EMNLP 2023 · 315 citations
Related papers
- Adaptive Retrieval Without Self-Knowledge? Bringing Uncertainty Back HomeViktor Moskvoretskii, Maria Marina, Mikhail Salnikov, Nikolay Ivanov et al.ACL 2025 · 22 citations
- SeaKR: Self-aware Knowledge Retrieval for Adaptive Retrieval Augmented GenerationZijun Yao, Weijian Qi, Liangming Pan, Shulin Cao et al.ACL 2025
- From Conflict to Consensus: Boosting Medical Reasoning via Multi-Round Agentic RAGWenhao Wu, Zhentao Tang, Yafu Li, Shixiong Kai et al.ICML 2026
- SGIC: A Self-Guided Iterative Calibration Framework for RAGGuanhua Chen, Yutong Yao, Lidia S. Chao, Xuebo Liu et al.ACL 2025
- NeocorRAG: Less Irrelevant Information, More Explicit Evidence, and More Effective Recall via Evidence ChainsShiyao Peng, Qianhe Zheng, Zhuodi Hao, Zichen Tang et al.WWW 2026
