Frustratingly Simple Retrieval Improves Challenging, Reasoning-Intensive Benchmarks
Xinxi Lyu, Michael Duan, Rulin Shao, Pang Wei Koh, Sewon Min
Abstract
Retrieval-augmented Generation (RAG) has primarily been studied in limited settings, such as factoid question answering; more challenging, reasoning-intensive benchmarks have seen limited success from minimal RAG. In this work, we challenge this prevailing view on established, reasoning-intensive benchmarks: MMLU, MMLU Pro, AGI Eval, GPQA, and MATH. We identify a key missing component in prior work: a usable, web-scale datastore aligned with the breadth of pretraining data. To this end, we introduce COMPACTDS: a diverse, high-quality, web-scale datastore that achieves high retrieval accuracy and subsecond latency on a single-node. The key insights are (1) most web content can be filtered out without sacrificing coverage, and a compact, high-quality subset is sufficient; and (2) combining in-memory approximate nearest neighbor (ANN) retrieval and ondisk exact search balances speed and recall. Using COMPACTDS, we show that a minimal RAG pipeline achieves consistent accuracy improvements across all benchmarks and model sizes (8B-70B), with relative gains of 10% on MMLU, 33% on MMLU Pro, 14% on GPQA, and 19% on MATH. No single data source suffices alone, highlighting the importance of diversity of sources (web crawls, curated math, academic papers, educational text). Finally, we show that our carefully designed in-house datastore matches or outperforms web search engines such as Google Search, as well as recently proposed, complex agent-based RAG systems-all while maintaining simplicity, reproducibility, and self-containment. We release COMPACTDS and our retrieval pipeline, supporting future research exploring retrieval-based AI systems. * Equal contribution. Preprint. Under review. Common Crawl, e.g., five billion tokens [6, 34, 4 ]. These efforts were still evaluated on perplexity or Wikipedia-based benchmarks (except for [6, 8] on MMLU, which we compare against). We argue that prior datastores are either too narrow or small to be broadly effective, or not practically usable, e.g., MassiveDS [8] requires over 12TB of RAM to avoid multi-minute latency, making deployment infeasible in typical academic settings without distributed infrastructure. This work directly addresses these issues, proposing a datastore that is large and broad in coverage, yet compact enough to enable subsecond latency in a single-node deployment. Agentic RAG. Recently, agentic RAG, which iteratively issues search queries, retrieves information, and reasons over results to perform reasoning-intensive tasks, has emerged as an active area of research. These approaches can be broadly divided into two categories: (1) prompt-based methods that do not require training [15, 16] , and (2) training-based methods that fine-tune a reasoning LM to use search, typically via reinforcement learning [17, 18, 19, 20] . Much of this work uses web search engines, which are costly, hard to reproduce, and unstable, making them unsuitable for training, as also noted by [20] . Consequently, most training-based work uses an in-house Wikipedia datastore and only evaluate on Wikipedia-based benchmarks. Instead of optimizing for agentic RAG, our work focuses on minimal RAG, which is a fundamental building block of any retrieval-based AI systems that can be easily integrated. This agentic RAG literature, however, highlights an emerging need for high-quality, general-purpose in-house datastores, particularly to enhance reproducibility, improve stability, and ensure cost efficiency. Method Two key ideas enable a high-quality, high-coverage retrieval datastore: data sources that match the breadth of pretraining corpora while filtering out low-quality web text ( §3.1), and approximate nearest neighbor (ANN) search followed by exact search ( §3.2). We discuss each component, then describe how an LLM is augmented with this retrieval ( §3.3). COMPACTDS Data Sources To match the breadth of pretraining corpora while achieving high quality and diversity, we strategically construct COMPACTDS with the following data sources: Web Crawl. To ensure wide coverage, we start with Common Crawl, which is widely used for pre-training and also constitutes 70% of MASSIVEDS [8] . However, we hypothesize that much of it is low-quality and unnecessary for retrieval. Therefore, we construct a compact, high-quality subset-High-quality CC-using a series of filtering steps. We take the union of C4 [35], a small curated subset, and DCLM-Baseline [36] , which has undergone extensive manual and model-based filtering. We further filter DCLM-baseline using the FineWeb-Edu classifier [33] with a threshold of 4.0, which filters text based on its educational value. Overall, this process reduces the size of
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e03e771b-3262-4967-9d81-8cdb05f2f291Cited by top-tier papers3
- Pretraining with hierarchical memories: separating long-tail and common knowledgeHadi Pouransari, David Grangier, C Thomas, Michael Kirchhof et al.ICLR 2026 · 11 citations
- Reusing Pre-Training Data at Test Time is a Compute MultiplierAlex Fang, Thomas Voice, Ruoming Pang, Ludwig Schmidt et al.ICLR 2026 · 4 citations
- Demystifying Scientific Problem-Solving in LLMs by Probing Knowledge and ReasoningAlan Li, Yixin Liu, Arpan Sarkar, Doug Downey et al.ICML 2026 · 4 citations
Builds on15
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat et al.ICML 2020 · 2,937 citations
Related papers
- The Power of Noise: Redefining Retrieval for RAG SystemsFlorin Cuconasu, Giovanni Trappolini, Federico Siciliano, Simone Filice et al.SIGIR 2024 · 212 citations
- In-Storage Acceleration of Retrieval Augmented Generation as a ServiceRohan Mahapatra, Harsha Santhanam, Christopher Priebe, Hanyang Xu et al.ISCA 2025 · 9 citations
- DReX: Accurate and Scalable Dense Retrieval Acceleration via Algorithmic-Hardware CodesignDerrick Quinn, E. Ezgi Yücel, Martin Prammer, Zhenxing Fan et al.ISCA 2025 · 10 citations
- The Retrieval Bottleneck: Scaling Laws for Reinforcement Learning in RAGShu Zhou, Jinman Leng, Yufei Song, Xin Wang et al.ACL 2026
- Accelerating Retrieval-Augmented GenerationDerrick Quinn, Mohammad Nouri, Neel Patel, John Salihu et al.ASPLOS 2025 · 37 citations
