Top-H Decoding: Adapting the Creativity and Coherence with Bounded Entropy in Text Generation
Erfan Baghaei Potraghloo, Seyedarmin Azizi, Souvik Kundu, Massoud Pedram
Abstract
Large language models (LLMs), despite their impressive performance across a wide range of tasks, often struggle to balance two competing objectives in open-ended text generation: fostering diversity and creativity while preserving logical coherence. Existing truncated sampling techniques, including temperature scaling, top-$p$ (nucleus) sampling, and min-$p$ sampling, aim to manage this trade-off. However, they exhibit limitations, particularly in the effective incorporation of the confidence of the model into the corresponding sampling strategy. For example, min-$p$ sampling relies on a single top token as a heuristic for confidence, eventually underutilizing the information of the probability distribution. Toward effective incorporation of the confidence of the model, in this paper, we present top-H decoding. We first establish the theoretical foundation of the interplay between creativity and coherence in truncated sampling by formulating an entropy-constrained minimum divergence problem. We then prove this minimization problem to be equivalent to an entropy-constrained mass maximization (ECMM) problem, which is NP-hard. Finally, we present top-H decoding, a computationally efficient greedy algorithm to solve the ECMM problem. Extensive empirical evaluations demonstrate that top-H outperforms the state-of-the-art (SoTA) alternative of min-$p$ sampling by up to 25.63% on creative writing benchmarks, while maintaining robustness on question-answering datasets such as GPQA, GSM8K, and MT-Bench. Additionally, an LLM-as-judge evaluation confirms that top-H indeed produces coherent outputs even at higher temperatures, where creativity is especially critical. In summary, top-H advances SoTA in open-ended text generation and can be easily integrated into creative writing applications. The code is available at https://github.com/ErfanBaghaei/Top-H-Decoding.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 429ea672-c711-40fe-afd2-603ac6472c0eCited by top-tier papers1
Ask how each one uses itBuilds on4
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- CoAuthor: Designing a Human-AI Collaborative Writing Dataset for Exploring Language Model CapabilitiesMina Lee, Percy Liang, Qian YangCHI 2022 · 340 citations
- Mirostat: a Neural Text decoding Algorithm that directly controls perplexitySourya Basu, Govardana Sachitanandam Ramachandran, Nitish Shirish Keskar, Lav R. VarshneyICLR 2021 · 13 citations
- Turning Up the Heat: Min-p Sampling for Creative and Coherent LLM OutputsNguyen Nhat Minh, Andrew Baker, Clement Neo, Allen G. Roush et al.ICLR 2025
Related papers
- Geometry-Aware Decoding with Wasserstein-Regularized Truncation and Mass Penalties for Large Language ModelsArash Gholamidavoodi, Navid Rezazadeh, Seyed Davoudi, Pouya PezeshkpourICML 2026 · 2 citations
- Min-k Sampling: Decoupling Truncation from Temperature Scaling via Relative Logit DynamicsYuanhao Ding, Meimingwei Li, Esteban Garces Arias, Matthias Aßenmacher et al.ACL 2026 · 4 citations
- p-less Sampling: A Robust Hyperparameter-Free Approach for LLM DecodingRunyan Tan, Shuang Wu, Phillip HowardICLR 2026 · 2 citations
- Sample Smart, Not Hard: Correctness-First Decoding for Better Reasoning in LLMsXueyan Li, Guinan Su, Mrinmaya Sachan, Jonas GeipingICLR 2026 · 5 citations
- Closing the Curious Case of Neural Text DegenerationMatthew Finlayson, John Hewitt, Alexander Koller, Swabha Swayamdipta et al.ICLR 2024 · 31 citations
