Evaluating n-Gram Novelty of Language Models Using Rusty-DAWG
William Merrill, Noah A. Smith, Yanai Elazar
摘要
How novel are texts generated by language models (LMs) relative to their training corpora? In this work, we investigate the extent to which modern LMs generate n-grams from their training data, evaluating both (i) the probability LMs assign to complete training n-grams and (ii) n-novelty, the proportion of n-grams generated by an LM that did not appear in the training data (for arbitrarily large n). To enable arbitrary-length n-gram search over a corpus in constant time w.r.t. corpus size, we develop RUSTY-DAWG, a novel search tool inspired by indexing of genomic data. We compare the novelty of LM-generated text to humanwritten text and explore factors that affect generation novelty, focusing on the Pythia models. We find that, for n > 4, LM-generated text is less novel than human-written text, though it is more novel for smaller n. Larger LMs and more constrained decoding strategies both decrease novelty. Finally, we show that LMs complete n-grams with lower loss if they are more frequent in the training data. Overall, our results reveal factors influencing the novelty of LMgenerated text, and we release RUSTY-DAWG to facilitate further pretraining data research. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Death of the Novel(ty): Beyond N-Gram Novelty as a Metric for Textual CreativityArkadiy Saakyan, Najoung Kim, Smaranda Muresan, Tuhin ChakrabartyICLR 2026 · 被引用 6 次
- Measuring LLM Novelty As The Frontier Of Original And High-Quality OutputVishakh Padmakumar, Chen Yueh-Han, Jane Pan, Valerie Chen 等ICLR 2026 · 被引用 6 次
- Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining DataXinyi Wang, Antonis Antoniades, Yanai Elazar, Alfonso Amayuelas 等ICLR 2025 · 被引用 4 次
- Detection and Measurement of Syntactic Templates in Generated TextChantal Shaib, Yanai Elazar, Junyi Jessy Li, Byron C. WallaceEMNLP 2024 · 被引用 3 次
- Infini-gram mini: Exact n-gram Search at the Internet Scale with FM-IndexHao Xu, Jiacheng Liu, Yejin Choi, Noah A. Smith 等EMNLP 2025 · 被引用 1 次
它引用的顶会 Paper14
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
- Generalization through Memorization: Nearest Neighbor Language ModelsUrvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer 等ICLR 2020 · 被引用 1,038 次
- Deduplicating Training Data Makes Language Models BetterKatherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang 等ACL 2022 · 被引用 844 次
相关 Paper
- DNA-GPT: Divergent N-Gram Analysis for Training-Free Detection of GPT-Generated TextXianjun Yang, Wei Cheng, Yue Wu, Linda Ruth Petzold 等ICLR 2024 · 被引用 173 次
- PolyPythias: Stability and Outliers across Fifty Language Model Pre-Training RunsOskar van der Wal, Pietro Lesci, Max Müller-Eberstein, Naomi Saphra 等ICLR 2025
- Comparing LLM-generated and human-authored news text using formal syntactic theoryOlga Zamaraeva, Dan Flickinger, Francis Bond, Carlos Gómez-RodríguezACL 2025 · 被引用 8 次
- A Tale of Two Structures: Do LLMs Capture the Fractal Complexity of Language?Ibrahim Alabdulmohsin, Andreas Peter SteinerICML 2025
- Exploring Training and Inference Scaling Laws in Generative RetrievalHongru Cai, Yongqi Li, Ruifeng Yuan, Wenjie Wang 等SIGIR 2025 · 被引用 1 次
