Top-nσ: Eliminating Noise in Logit Space for Robust Token Sampling of LLM
Chenxia Tang, Jianchun Liu, Hongli Xu, Liusheng Huang
Abstract
Large language models (LLMs) rely heavily on sampling methods to generate diverse and highquality text. While existing sampling methods like top-p and min-p have identified the detrimental effects of low-probability tails in LLMs' outputs, they still fail to effectively distinguish between diversity and noise. This limitation stems from their reliance on probability-based metrics that are inherently sensitive to temperature scaling. Through empirical and theoretical analysis, we make two key discoveries: (1) the pre-softmax logits exhibit a clear statistical separation between informative tokens and noise, and (2) we prove the mathematical equivalence of min-p and top-(1-p) under uniform distribution over logits. These findings motivate the design of top-nσ, a novel sampling method that identifies informative tokens by eliminating noise directly in logit space. Unlike existing methods that become unstable at high temperatures, top-nσ achieves temperature-invariant token selection while preserving output diversity. Extensive experiments across reasoning and creative writing tasks demonstrate that our method consistently outperforms existing approaches, with particularly significant improvements in high-temperature settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 05f60341-1b5e-4e7f-bc06-65006880b00aCited by top-tier papers3
- Min-k Sampling: Decoupling Truncation from Temperature Scaling via Relative Logit DynamicsYuanhao Ding, Meimingwei Li, Esteban Garces Arias, Matthias Aßenmacher et al.ACL 2026 · 4 citations
- Verifiable LLM-Generated Text Detection via Projected Semantic-Structural DistributionsRuochong Xiong, Qien Li, Wangwang Lian, Yulong Wan et al.ACL 2026
- Beyond Temperature: Hyperfitting as a Late-Stage Geometric ExpansionMeimingwei Li, Yuanhao Ding, Esteban Garces Arias, Christian HeumannICML 2026
Builds on6
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer et al.NeurIPS 2022 · 2,039 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Mirostat: a Neural Text decoding Algorithm that directly controls perplexitySourya Basu, Govardana Sachitanandam Ramachandran, Nitish Shirish Keskar, Lav R. VarshneyICLR 2021 · 13 citations
Related papers
- Turning Up the Heat: Min-p Sampling for Creative and Coherent LLM OutputsNguyen Nhat Minh, Andrew Baker, Clement Neo, Allen G. Roush et al.ICLR 2025
- p-less Sampling: A Robust Hyperparameter-Free Approach for LLM DecodingRunyan Tan, Shuang Wu, Phillip HowardICLR 2026 · 2 citations
- Closing the Curious Case of Neural Text DegenerationMatthew Finlayson, John Hewitt, Alexander Koller, Swabha Swayamdipta et al.ICLR 2024 · 31 citations
- Top-H Decoding: Adapting the Creativity and Coherence with Bounded Entropy in Text GenerationErfan Baghaei Potraghloo, Seyedarmin Azizi, Souvik Kundu, Massoud PedramNeurIPS 2025 · 13 citations
- LLM-Oriented Token-Adaptive Knowledge DistillationXurong Xie, Zhucun Xue, Jiafu Wu, Jian Li et al.AAAI 2026
