Long Horizon Temperature Scaling
Andy Shih, Dorsa Sadigh, Stefano Ermon
Abstract
Temperature scaling is a popular technique for tuning the sharpness of a model distribution. It is used extensively for sampling likely generations and calibrating model uncertainty, and even features as a controllable parameter to many large language models in deployment. However, autoregressive models rely on myopic temperature scaling that greedily optimizes the next token. To address this, we propose Long Horizon Temperature Scaling (LHTS), a novel approach for sampling from temperature-scaled joint distributions. LHTS is compatible with all likelihoodbased models, and optimizes for the long horizon likelihood of samples. We derive a temperaturedependent LHTS objective, and show that finetuning a model on a range of temperatures produces a single model capable of generation with a controllable long horizon temperature parameter. We experiment with LHTS on image diffusion models and character/language autoregressive models, demonstrating advantages over myopic temperature scaling in likelihood and sample quality, and showing improvements in accuracy on a multiple choice analogy task by 10%. Our code is available at https://github.com/AndyShih12/ LongHorizonTemperatureScaling .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers15
- MEGA: Multilingual Evaluation of Generative AIKabir Ahuja, Harshita Diddee, Rishav Hada, Millicent Ochieng et al.EMNLP 2023 · 91 citations
- Amortizing intractable inference in large language modelsEdward J. Hu, Moksh Jain, Eric Elmoznino, Younesse Kaddar et al.ICLR 2024 · 91 citations
- Probabilistic Inference in Language Models via Twisted Sequential Monte CarloStephen Zhao, Rob Brekelmans, Alireza Makhzani, Roger Baker GrosseICML 2024 · 61 citations
- Verifying Chain-of-Thought Reasoning via Its Computational GraphZheng Zhao, Yeskendir Koishekenov, Xianjun Yang, Naila Murray et al.ICLR 2026 · 25 citations
- Post-Hoc Reversal: Are We Selecting Models Prematurely?Rishabh Ranjan, Saurabh Garg, Mrigank Raman, Carlos Guestrin et al.NeurIPS 2024 · 8 citations
Builds on7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 5,234 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- NVAE: A Deep Hierarchical Variational AutoencoderArash Vahdat, Jan KautzNeurIPS 2020 · 1,141 citations
Related papers
- Temperature Scaling in Discrete Sequence (Language) ModelsHannah Scheufele, Peter Blohm, Vikas GargICML 2026
- Towards Better & Faster Autoregressive Image Generation: From the Perspective of EntropyXiaoxiao Ma, Feng Zhao, Pengyang Ling, Haibo Qiu et al.NeurIPS 2025 · 12 citations
- Improving Diversity in Language Models: When Temperature Fails, Change the LossAlexandre Verine, Florian Le Bronnec, Kunhao Zheng, Alexandre Allauzen et al.ICML 2025
- Large Language Models to Diffusion FinetuningEdoardo Cetin, Tianyu Zhao, Yujin TangICML 2025
- Reward-Weighted Sampling: Enhancing Non-Autoregressive Characteristics in Masked Diffusion LLMsDaehoon Gwak, Minseo Jung, Junwoo Park, Minho Park et al.EMNLP 2025
