Evaluating n-Gram Novelty of Language Models Using Rusty-DAWG
William Merrill, Noah A. Smith, Yanai Elazar
Abstract
How novel are texts generated by language models (LMs) relative to their training corpora? In this work, we investigate the extent to which modern LMs generate n-grams from their training data, evaluating both (i) the probability LMs assign to complete training n-grams and (ii) n-novelty, the proportion of n-grams generated by an LM that did not appear in the training data (for arbitrarily large n). To enable arbitrary-length n-gram search over a corpus in constant time w.r.t. corpus size, we develop RUSTY-DAWG, a novel search tool inspired by indexing of genomic data. We compare the novelty of LM-generated text to humanwritten text and explore factors that affect generation novelty, focusing on the Pythia models. We find that, for n > 4, LM-generated text is less novel than human-written text, though it is more novel for smaller n. Larger LMs and more constrained decoding strategies both decrease novelty. Finally, we show that LMs complete n-grams with lower loss if they are more frequent in the training data. Overall, our results reveal factors influencing the novelty of LMgenerated text, and we release RUSTY-DAWG to facilitate further pretraining data research. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e9a3fdde-a0d9-430a-b922-e07a5b17b120Cited by top-tier papers8
- Death of the Novel(ty): Beyond N-Gram Novelty as a Metric for Textual CreativityArkadiy Saakyan, Najoung Kim, Smaranda Muresan, Tuhin ChakrabartyICLR 2026 · 6 citations
- Measuring LLM Novelty As The Frontier Of Original And High-Quality OutputVishakh Padmakumar, Chen Yueh-Han, Jane Pan, Valerie Chen et al.ICLR 2026 · 6 citations
- Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining DataXinyi Wang, Antonis Antoniades, Yanai Elazar, Alfonso Amayuelas et al.ICLR 2025 · 4 citations
- Detection and Measurement of Syntactic Templates in Generated TextChantal Shaib, Yanai Elazar, Junyi Jessy Li, Byron C. WallaceEMNLP 2024 · 3 citations
- Infini-gram mini: Exact n-gram Search at the Internet Scale with FM-IndexHao Xu, Jiacheng Liu, Yejin Choi, Noah A. Smith et al.EMNLP 2025 · 1 citation
Builds on14
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- Generalization through Memorization: Nearest Neighbor Language ModelsUrvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer et al.ICLR 2020 · 1,038 citations
- Deduplicating Training Data Makes Language Models BetterKatherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang et al.ACL 2022 · 844 citations
Related papers
- DNA-GPT: Divergent N-Gram Analysis for Training-Free Detection of GPT-Generated TextXianjun Yang, Wei Cheng, Yue Wu, Linda Ruth Petzold et al.ICLR 2024 · 173 citations
- PolyPythias: Stability and Outliers across Fifty Language Model Pre-Training RunsOskar van der Wal, Pietro Lesci, Max Müller-Eberstein, Naomi Saphra et al.ICLR 2025
- Comparing LLM-generated and human-authored news text using formal syntactic theoryOlga Zamaraeva, Dan Flickinger, Francis Bond, Carlos Gómez-RodríguezACL 2025 · 8 citations
- A Tale of Two Structures: Do LLMs Capture the Fractal Complexity of Language?Ibrahim Alabdulmohsin, Andreas Peter SteinerICML 2025
- Exploring Training and Inference Scaling Laws in Generative RetrievalHongru Cai, Yongqi Li, Ruifeng Yuan, Wenjie Wang et al.SIGIR 2025 · 1 citation
