Infini-gram mini: Exact n-gram Search at the Internet Scale with FM-Index
Hao Xu, Jiacheng Liu, Yejin Choi, Noah A. Smith, Hannaneh Hajishirzi
Abstract
Language models are trained mainly on massive text data from the Internet, and it becomes increasingly important to understand this data source. Exact-match search engines enable searching in large text corpora -counting string appearances and retrieving the enclosing documents -yet the high storage overhead hinders their application on Internet-scale data. We present INFINI-GRAM MINI, an efficient and scalable system that can make petabyte-level text corpora searchable. Based on the FMindex data structure (Ferragina and Manzini, 2000), which simultaneously indexes and compresses text, our system creates indexes with size only 44% of the corpus. INFINI-GRAM MINI greatly improves upon the best existing implementation of FM-index in terms of indexing speed (18×) and memory use during both indexing (3.2× reduction) and querying (down to a negligible amount). We index 83TB of Internet text in 99 days with a single CPU node with 128 vCPUs (or 19 hours if using 137 such nodes). We show one important use case of INFINI-GRAM MINI in a large-scale analysis of benchmark contamination. We find several core LM evaluation benchmarks to be heavily contaminated in Internet crawls (up to 74.2% in GSM8K), which could lead to overestimating the capabilities of language models if trained on such data. We host a benchmark contamination bulletin to share the contamination rate of many core and community-contributed benchmarks. We also release a web interface and an API endpoint to serve general search queries on INFINI-GRAM MINI indexes. Project Home infini-gram-mini.io Web Interface infini-gram-mini.io/demo API Endpoint api.infini-gram-mini.io Documentation infini-gram-mini.io/docs Source Code infini-gram-mini.io/code Contam Bulletin infini-gram-mini.io/bulletin
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c1d21f04-7a8d-44e2-8375-a948cb706162Cited by top-tier papers5
- Speculative Speculative DecodingTanishq Kumar, Tri Dao, Avner MayICLR 2026 · 15 citations
- Death of the Novel(ty): Beyond N-Gram Novelty as a Metric for Textual CreativityArkadiy Saakyan, Najoung Kim, Smaranda Muresan, Tuhin ChakrabartyICLR 2026 · 6 citations
- BenchMarker: An Education-Inspired Toolkit for Highlighting Flaws in Multiple-Choice BenchmarksNishant Balepur, Bhavya Rajasekaran, Hyunjin Jane Oh, Michael Xie et al.ACL 2026 · 1 citation
- SoftMatcha 2: A Fast and Soft Pattern Matcher for Trillion-Scale CorporaMasataka Yoneda, Yusuke Matsushita, Go Kamoda, Kohei Suenaga et al.ICML 2026
- Characterizing and Evaluating Working Emotion Vocabularies in Multilingual Large Language ModelsNicholas Deas, Iván Ernesto Pérez Mejía, Ellie Yang, Kathleen McKeownACL 2026
Builds on9
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- Deduplicating Training Data Makes Language Models BetterKatherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang et al.ACL 2022 · 844 citations
Related papers
- What's In My Big Data?Yanai Elazar, Akshita Bhagia, Ian Magnusson, Abhilasha Ravichander et al.ICLR 2024 · 135 citations
- Dripper: Token-Efficient Main HTML Extraction with a Lightweight LMMengjie Liu, Jiahui Peng, Wenchang Ning, Pei Chu et al.KDD 2026 · 10 citations
- MiniRAG: A Lightweight RAG system with Small Language ModelsTianyu Fan, Jingyuan Wang, Xubin Ren, Chao HuangACL 2026
- LogCloud: Fast Search of Compressed Logs on Object StorageZiheng Wang, Junyu Wei, Alex Aiken, Guangyan Zhang et al.VLDB 2025 · 1 citation
- Text-to-ES Bench: A Comprehensive Benchmark for Converting Natural Language to Elasticsearch QueryDongge Xue, Zhili Pu, Zhentao Xia, Hongli Sun et al.ACL 2025
