Top-nσ: Eliminating Noise in Logit Space for Robust Token Sampling of LLM
Chenxia Tang, Jianchun Liu, Hongli Xu, Liusheng Huang
摘要
Large language models (LLMs) rely heavily on sampling methods to generate diverse and highquality text. While existing sampling methods like top-p and min-p have identified the detrimental effects of low-probability tails in LLMs' outputs, they still fail to effectively distinguish between diversity and noise. This limitation stems from their reliance on probability-based metrics that are inherently sensitive to temperature scaling. Through empirical and theoretical analysis, we make two key discoveries: (1) the pre-softmax logits exhibit a clear statistical separation between informative tokens and noise, and (2) we prove the mathematical equivalence of min-p and top-(1-p) under uniform distribution over logits. These findings motivate the design of top-nσ, a novel sampling method that identifies informative tokens by eliminating noise directly in logit space. Unlike existing methods that become unstable at high temperatures, top-nσ achieves temperature-invariant token selection while preserving output diversity. Extensive experiments across reasoning and creative writing tasks demonstrate that our method consistently outperforms existing approaches, with particularly significant improvements in high-temperature settings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Min-k Sampling: Decoupling Truncation from Temperature Scaling via Relative Logit DynamicsYuanhao Ding, Meimingwei Li, Esteban Garces Arias, Matthias Aßenmacher 等ACL 2026 · 被引用 4 次
- Verifiable LLM-Generated Text Detection via Projected Semantic-Structural DistributionsRuochong Xiong, Qien Li, Wangwang Lian, Yulong Wan 等ACL 2026
- Beyond Temperature: Hyperfitting as a Late-Stage Geometric ExpansionMeimingwei Li, Yuanhao Ding, Esteban Garces Arias, Christian HeumannICML 2026
它引用的顶会 Paper6
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer 等NeurIPS 2022 · 被引用 2,039 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Mirostat: a Neural Text decoding Algorithm that directly controls perplexitySourya Basu, Govardana Sachitanandam Ramachandran, Nitish Shirish Keskar, Lav R. VarshneyICLR 2021 · 被引用 13 次
相关 Paper
- Turning Up the Heat: Min-p Sampling for Creative and Coherent LLM OutputsNguyen Nhat Minh, Andrew Baker, Clement Neo, Allen G. Roush 等ICLR 2025
- p-less Sampling: A Robust Hyperparameter-Free Approach for LLM DecodingRunyan Tan, Shuang Wu, Phillip HowardICLR 2026 · 被引用 2 次
- Closing the Curious Case of Neural Text DegenerationMatthew Finlayson, John Hewitt, Alexander Koller, Swabha Swayamdipta 等ICLR 2024 · 被引用 31 次
- Top-H Decoding: Adapting the Creativity and Coherence with Bounded Entropy in Text GenerationErfan Baghaei Potraghloo, Seyedarmin Azizi, Souvik Kundu, Massoud PedramNeurIPS 2025 · 被引用 13 次
- LLM-Oriented Token-Adaptive Knowledge DistillationXurong Xie, Zhucun Xue, Jiafu Wu, Jian Li 等AAAI 2026
