Closing the Curious Case of Neural Text Degeneration
Matthew Finlayson, John Hewitt, Alexander Koller, Swabha Swayamdipta, Ashish Sabharwal
摘要
Despite their ubiquity in language generation, it remains unknown why truncation sampling heuristics like nucleus sampling are so effective. We provide a theoretical explanation for the effectiveness of the truncation sampling by proving that truncation methods that discard tokens below some probability threshold (the most common type of truncation) can guarantee that all sampled tokens have nonzero true probability. However, thresholds are a coarse heuristic, and necessarily discard some tokens with nonzero true probability as well. In pursuit of a more precise sampling strategy, we show that we can leverage a known source of model errors, the softmax bottleneck, to prove that certain tokens have nonzero true probability, without relying on a threshold. Based on our findings, we develop an experimental truncation strategy and the present pilot studies demonstrating the promise of this type of algorithm. Our evaluations show that our method outperforms its threshold-based counterparts under automatic and human evaluation metrics for low-entropy (i.e., close to greedy) open-ended text generation. Our theoretical findings and pilot experiments provide both insight into why truncation sampling works, and make progress toward more expressive sampling algorithms that better surface the generative capabilities of large language models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Towards Understanding Subliminal Learning: When and How Hidden Biases TransferSimon Schrodi, Elias Kempf, Fazl Barez, Thomas BroxICLR 2026 · 被引用 30 次
- Improving Open-Ended Text Generation via Adaptive DecodingWenhong Zhu, Hongkun Hao, Zhiwei He, Yiming Ai 等ICML 2024 · 被引用 21 次
- Foundations of Top-k Decoding for Language ModelsGeorgy Noarov, Soham Mallick, Tao Wang, Sunay Joshi 等NeurIPS 2025 · 被引用 14 次
- Better Language Model Inversion by Compactly Representing Next-Token DistributionsMurtaza Nazir, Matthew Finlayson, John X. Morris, Xiang Ren 等NeurIPS 2025 · 被引用 13 次
- InvisibleInk: High-Utility and Low-Cost Text Generation with Differential PrivacyVishnu Vinod, Krishna Pillutla, Abhradeep Guha ThakurtaNeurIPS 2025 · 被引用 12 次
它引用的顶会 Paper8
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
- MAUVE: Measuring the Gap Between Neural Text and Human Text using Divergence FrontiersKrishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun 等NeurIPS 2021 · 被引用 606 次
- Contrastive Decoding: Open-ended Text Generation as OptimizationXiang Lisa Li, Ari Holtzman, Daniel Fried, Percy Liang 等ACL 2023 · 被引用 78 次
- RankGen: Improving Text Generation with Large Ranking ModelsKalpesh Krishna, Yapei Chang, John Wieting, Mohit IyyerEMNLP 2022 · 被引用 27 次
相关 Paper
- Turning Up the Heat: Min-p Sampling for Creative and Coherent LLM OutputsNguyen Nhat Minh, Andrew Baker, Clement Neo, Allen G. Roush 等ICLR 2025
- Decoding Game: On Minimax Optimality of Heuristic Text Generation StrategiesSijin Chen, Omar Hagrass, Jason Matthew KlusowskiICLR 2025
- Automatic Detection of Generated Text is Easiest when Humans are FooledDaphne Ippolito, Daniel Duckworth, Chris Callison-Burch, Douglas EckACL 2020 · 被引用 21 次
- Top-nσ: Eliminating Noise in Logit Space for Robust Token Sampling of LLMChenxia Tang, Jianchun Liu, Hongli Xu, Liusheng HuangACL 2025 · 被引用 9 次
- Entropy-informed Decoding: Adaptive Information-Driven BranchingBenjamin Patrick Evans, Sumitra Ganesh, Leo ArdonICML 2026
