Inference Scaling for Long-Context Retrieval Augmented Generation
Zhenrui Yue, Honglei Zhuang, Aijun Bai, Kai Hui, Rolf Jagerman, Hansi Zeng, Zhen Qin, Dong Wang, Xuanhui Wang, Michael Bendersky
Abstract
The scaling of inference computation has unlocked the potential of long-context large language models (LLMs) across diverse settings. For knowledge-intensive tasks, the increased compute is often allocated to incorporate more external knowledge. However, without effectively utilizing such knowledge, solely expanding context does not always enhance performance. In this work, we investigate inference scaling for retrieval augmented generation (RAG), exploring the combination of multiple strategies beyond simply increasing the quantity of knowledge, including in-context learning and iterative prompting. These strategies provide additional flexibility to scale test-time computation (e.g., by increasing retrieved documents or generation steps), thereby enhancing LLMs' ability to effectively acquire and utilize contextual information. We address two key questions: (1) How does RAG performance benefit from the scaling of inference computation when optimally configured? (2) Can we predict the optimal test-time compute allocation for a given budget by modeling the relationship between RAG performance and inference parameters? Our observations reveal that increasing inference computation leads to nearly linear gains in RAG performance when optimally allocated, a relationship we describe as the inference scaling laws for RAG. Building on this, we further develop the computation allocation model to estimate RAG performance across different inference configurations. The model predicts optimal inference parameters under various computation constraints, which align closely with the experimental results. By applying these optimal configurations, we demonstrate that scaling inference compute on long-context LLMs achieves up to 58.9% gains on benchmark datasets compared to standard RAG.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers34
- A Survey of Large Language Model-Based Search AgentsYunjia Xi, Jianghao Lin, Yongzhao Xiao, Zheli Zhou et al.ACL 2026 · 1,216 citations
- MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon AgentsZijian Zhou, Ao Qu, Zhaoxuan Wu, Sunghwan Kim et al.ICLR 2026 · 223 citations
- Process vs. Outcome Reward: Which is Better for Agentic RAG Reinforcement LearningWenlin Zhang, Xiangyang Li, Kuicai Dong, Yichao Wang et al.NeurIPS 2025 · 85 citations
- Chain-of-Retrieval Augmented GenerationLiang Wang, Haonan Chen, Nan Yang, Xiaolong Huang et al.NeurIPS 2025 · 59 citations
- DeepRAG: Thinking to Retrieve Step by Step for Large Language ModelsXinyan Guan, Jiali Zeng, Fandong Meng, Chunlei Xin et al.ICLR 2026 · 30 citations
Builds on35
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
Related papers
- Inference Scaling Law for Retrieval Augmented GenerationShu Zhou, Yuxuan Ao, Yunyang Xuan, Xin Wang et al.AAAI 2026 · 1 citation
- Incentivizing Retrieval-Augmented Generation via Inner Adaptive Context SelectionChenxu Cui, Lin Shen, Haihui Fan, Sa Zhu et al.SIGIR 2026
- MemoRAG: Boosting Long Context Processing with Global Memory-Enhanced Retrieval AugmentationHongjin Qian, Zheng Liu, Peitian Zhang, Kelong Mao et al.WWW 2025 · 92 citations
- Superposition Prompting: Improving and Accelerating Retrieval-Augmented GenerationThomas Merth, Qichen Fu, Mohammad Rastegari, Mahyar NajibiICML 2024 · 14 citations
- Provence: efficient and robust context pruning for retrieval-augmented generationNadezhda Chirkova, Thibault Formal, Vassilina Nikoulina, Stéphane ClinchantICLR 2025 · 2 citations
