Foundations of Top-k Decoding for Language Models
Georgy Noarov, Soham Mallick, Tao Wang, Sunay Joshi, Yan Sun, Yangxinyu Xie, Mengxin Yu, Edgar Dobriban
摘要
Top-k decoding is a widely used method for sampling from LLMs: at each token, only the largest k next-token-probabilities are kept, and the next token is sampled after renormalizing them to sum to unity. Top-k and other sampling methods are motivated by the intuition that true next-token distributions are sparse, and the noisy LLM probabilities need to be truncated. However, to our knowledge, a precise theoretical motivation for the use of top-k decoding is missing. In this work, we develop a theoretical framework that both explains and generalizes top-k decoding. We view decoding at a fixed token as the recovery of a sparse probability distribution. We introduce Bregman decoders obtained by minimizing a separable Bregman divergence (for both the primal and dual cases) with a sparsity-inducing ℓ 0 -regularization; in particular, these decoders are adaptive in the sense that the sparsity parameter k is chosen depending on the underlying token distribution. Despite the combinatorial nature of the sparse Bregman objective, we show how to optimize it efficiently for a large class of divergences. We prove that (i) the optimal decoding strategies are greedy, and further that (ii) the objective is discretely convex in k, such that the optimal k can be identified in logarithmic time. We note that standard top-k decoding arises as a special case for the KL divergence, and construct new decoding strategies with substantially different behaviors (e.g., non-linearly up-weighting larger probabilities after renormalization).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper6
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Fast, Differentiable and Sparse Top-k: a Convex Analysis PerspectiveMichael Eli Sander, Joan Puigcerver, Josip Djolonga, Gabriel Peyré 等ICML 2023 · 被引用 35 次
- Closing the Curious Case of Neural Text DegenerationMatthew Finlayson, John Hewitt, Alexander Koller, Swabha Swayamdipta 等ICLR 2024 · 被引用 31 次
- Mirostat: a Neural Text decoding Algorithm that directly controls perplexitySourya Basu, Govardana Sachitanandam Ramachandran, Nitish Shirish Keskar, Lav R. VarshneyICLR 2021 · 被引用 13 次
- Decoding Game: On Minimax Optimality of Heuristic Text Generation StrategiesSijin Chen, Omar Hagrass, Jason Matthew KlusowskiICLR 2025
相关 Paper
- Min-k Sampling: Decoupling Truncation from Temperature Scaling via Relative Logit DynamicsYuanhao Ding, Meimingwei Li, Esteban Garces Arias, Matthias Aßenmacher 等ACL 2026 · 被引用 4 次
- On the Efficacy of Sampling AdaptersClara Meister, Tiago Pimentel, Luca Malagutti, Ethan Wilcox 等ACL 2023 · 被引用 3 次
- KLASS: KL-Guided Fast Inference in Masked Diffusion ModelsSeo Hyun Kim, Sunwoo Hong, Hojung Jung, Youngrok Park 等NeurIPS 2025 · 被引用 48 次
- Sample Smart, Not Hard: Correctness-First Decoding for Better Reasoning in LLMsXueyan Li, Guinan Su, Mrinmaya Sachan, Jonas GeipingICLR 2026 · 被引用 5 次
- Hot or Cold? Adaptive Temperature Sampling for Code Generation with Large Language ModelsYuqi Zhu, Jia Li, Ge Li, Yunfei Zhao 等AAAI 2024 · 被引用 68 次
