Efficient Context Selection for Long-Context QA: No Tuning, No Iteration, Just Adaptive-k
Chihiro Taguchi, Seiji Maekawa, Nikita Bhutani
Abstract
Retrieval-augmented generation (RAG) and long-context language models (LCLMs) both address context limitations of LLMs in opendomain question answering (QA). However, optimal external context to retrieve remains an open problem: fixing the retrieval size risks either wasting tokens or omitting key evidence. Existing adaptive methods like Self-RAG and SELF-ROUTE rely on iterative LLM prompting and perform well on factoid QA, but struggle with aggregation QA, where the optimal context size is both unknown and variable. We present Adaptive-k retrieval, a simple and effective single-pass method that adaptively selects the number of passages based on the distribution of the similarity scores between the query and the candidate passages. It does not require model fine-tuning, extra LLM inferences or changes to existing retriever-reader pipelines. On both factoid and aggregation QA benchmarks, Adaptive-k matches or outperforms fixed-k baselines while using up to 10× fewer tokens than full-context input, yet still retrieves 70% of relevant passages. It improves accuracy across five LCLMs and two embedding models, highlighting that dynamically adjusting context size leads to more efficient and accurate QA. 1 * Work done during internship at Megagon Labs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 890183eb-6b35-40b6-98c0-3dbf78ef0fd3Cited by top-tier papers1
Ask how each one uses itBuilds on4
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil et al.ICLR 2024 · 1,798 citations
- Searching for Best Practices in Retrieval-Augmented GenerationXiaohua Wang, Zhenghua Wang, Xuan Gao, Feiran Zhang et al.EMNLP 2024 · 75 citations
- Long-Context LLMs Meet RAG: Overcoming Challenges for Long Inputs in RAGBowen Jin, Jinsung Yoon, Jiawei Han, Sercan Ö. ArikICLR 2025
- Holistic Reasoning with Long-Context LMs: A Benchmark for Database Operations on Massive Textual DataSeiji Maekawa, Hayate Iso, Nikita BhutaniICLR 2025
Related papers
- SeaKR: Self-aware Knowledge Retrieval for Adaptive Retrieval Augmented GenerationZijun Yao, Weijian Qi, Liangming Pan, Shulin Cao et al.ACL 2025
- Adaptive Retrieval Without Self-Knowledge? Bringing Uncertainty Back HomeViktor Moskvoretskii, Maria Marina, Mikhail Salnikov, Nikolay Ivanov et al.ACL 2025 · 22 citations
- Q-RAG: Long Context Multi‑Step Retrieval via Value‑Based Embedder TrainingArtyom Y. Sorokin, Nazar Buzun, Alexander Anokhin, Egor Vedernikov et al.ICLR 2026 · 4 citations
- LaRA: Benchmarking Retrieval-Augmented Generation and Long-Context LLMs - No Silver Bullet for LC or RAG RoutingKuan Li, Liwen Zhang, Yong Jiang, Pengjun Xie et al.ICML 2025
- SmartChunk Retrieval: Query-Aware Chunk Compression with Planning for Efficient Document RAGXuechen Zhang, Koustava Goswami, Samet Oymak, Jiasi Chen et al.ICLR 2026 · 1 citation
