Infini-gram mini: Exact n-gram Search at the Internet Scale with FM-Index
Hao Xu, Jiacheng Liu, Yejin Choi, Noah A. Smith, Hannaneh Hajishirzi
摘要
Language models are trained mainly on massive text data from the Internet, and it becomes increasingly important to understand this data source. Exact-match search engines enable searching in large text corpora -counting string appearances and retrieving the enclosing documents -yet the high storage overhead hinders their application on Internet-scale data. We present INFINI-GRAM MINI, an efficient and scalable system that can make petabyte-level text corpora searchable. Based on the FMindex data structure (Ferragina and Manzini, 2000), which simultaneously indexes and compresses text, our system creates indexes with size only 44% of the corpus. INFINI-GRAM MINI greatly improves upon the best existing implementation of FM-index in terms of indexing speed (18×) and memory use during both indexing (3.2× reduction) and querying (down to a negligible amount). We index 83TB of Internet text in 99 days with a single CPU node with 128 vCPUs (or 19 hours if using 137 such nodes). We show one important use case of INFINI-GRAM MINI in a large-scale analysis of benchmark contamination. We find several core LM evaluation benchmarks to be heavily contaminated in Internet crawls (up to 74.2% in GSM8K), which could lead to overestimating the capabilities of language models if trained on such data. We host a benchmark contamination bulletin to share the contamination rate of many core and community-contributed benchmarks. We also release a web interface and an API endpoint to serve general search queries on INFINI-GRAM MINI indexes. Project Home infini-gram-mini.io Web Interface infini-gram-mini.io/demo API Endpoint api.infini-gram-mini.io Documentation infini-gram-mini.io/docs Source Code infini-gram-mini.io/code Contam Bulletin infini-gram-mini.io/bulletin
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Speculative Speculative DecodingTanishq Kumar, Tri Dao, Avner MayICLR 2026 · 被引用 15 次
- Death of the Novel(ty): Beyond N-Gram Novelty as a Metric for Textual CreativityArkadiy Saakyan, Najoung Kim, Smaranda Muresan, Tuhin ChakrabartyICLR 2026 · 被引用 6 次
- BenchMarker: An Education-Inspired Toolkit for Highlighting Flaws in Multiple-Choice BenchmarksNishant Balepur, Bhavya Rajasekaran, Hyunjin Jane Oh, Michael Xie 等ACL 2026 · 被引用 1 次
- SoftMatcha 2: A Fast and Soft Pattern Matcher for Trillion-Scale CorporaMasataka Yoneda, Yusuke Matsushita, Go Kamoda, Kohei Suenaga 等ICML 2026
- Characterizing and Evaluating Working Emotion Vocabularies in Multilingual Large Language ModelsNicholas Deas, Iván Ernesto Pérez Mejía, Ellie Yang, Kathleen McKeownACL 2026
它引用的顶会 Paper9
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 被引用 2,317 次
- Deduplicating Training Data Makes Language Models BetterKatherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang 等ACL 2022 · 被引用 844 次
相关 Paper
- What's In My Big Data?Yanai Elazar, Akshita Bhagia, Ian Magnusson, Abhilasha Ravichander 等ICLR 2024 · 被引用 135 次
- Dripper: Token-Efficient Main HTML Extraction with a Lightweight LMMengjie Liu, Jiahui Peng, Wenchang Ning, Pei Chu 等KDD 2026 · 被引用 10 次
- MiniRAG: A Lightweight RAG system with Small Language ModelsTianyu Fan, Jingyuan Wang, Xubin Ren, Chao HuangACL 2026
- LogCloud: Fast Search of Compressed Logs on Object StorageZiheng Wang, Junyu Wei, Alex Aiken, Guangyan Zhang 等VLDB 2025 · 被引用 1 次
- Text-to-ES Bench: A Comprehensive Benchmark for Converting Natural Language to Elasticsearch QueryDongge Xue, Zhili Pu, Zhentao Xia, Hongli Sun 等ACL 2025
