Optimizing Temperature for Language Models with Multi-Sample Inference
Weihua Du, Yiming Yang, Sean Welleck
Abstract
Multi-sample aggregation strategies, such as majority voting and best-of-N sampling, are widely used in contemporary large language models (LLMs) to enhance predictive accuracy across various tasks. A key challenge in this process is temperature selection, which significantly impacts model performance. Existing approaches either rely on a fixed default temperature or require labeled validation data for tuning, which are often scarce and difficult to obtain. This paper addresses the challenge of automatically identifying the (near)-optimal temperature for different LLMs using multi-sample aggregation strategies, without relying on task-specific validation data. We provide a comprehensive analysis of temperature's role in performance optimization, considering variations in model architectures, datasets, task types, model sizes, and predictive accuracy. Furthermore, we propose a novel entropy-based metric for automated temperature optimization, which consistently outperforms fixed-temperature baselines. Additionally, we incorporate a stochastic process model to enhance interpretability, offering deeper insights into the relationship between temperature and model performance. Our code is available at https://github.com/StigLidu/dualdistill . Optimizing Temperature for Language Models with Multi-Sample Inference ing how to optimize the sampling process to enhance LLM performance under different conditions, including variations in training datasets, task types, and model sizes. A crucial open question is how to tune temperature, a key hyperparameter that controls the smoothness of the system-learned distribution. Intuitively, increasing the temperature leads to a smoother distribution, enhancing the diversity of sampled outputs. However, excessively high temperatures can introduce many low-quality samples, making aggregation more challenging (Holtzman et al., 2019; Renze & Guven, 2024) . Conversely, lowering the temperature results in a highly concentrated distribution, reducing diversity and potentially omitting high-quality samples. Striking the right balance between over-sampling and under-sampling is therefore essential for optimizing LLM performance. A common practice in prior evaluations is to use the same temperature across all methods despite variations in training datasets, task types, model sizes, and aggregation strategies. This practice is clearly suboptimal. An alternative approach is to empirically tune the temperature using labeled validation data for each task, dataset, model size, and aggregation strategy (Zhang et al., 2024a; Dhuliawala et al., 2024) . However, such a process is tedious and time-consuming and heavily dependent on the availability of labeled validation data, limiting its applicability when such data are scarce. In this paper, we present the first systematic investigation of how temperature affects LLM performance under multisample aggregation strategies across various conditions. Furthermore, we propose a principled algorithmic solution for automated temperature optimization without requiring labeled validation data. Our key idea is as follows: 1. We use the confidence score of each model as a selfassessment measure.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3b5ddf5d-b442-4cac-a335-76fad6133d4bCited by top-tier papers11
- BroRL: Scaling Reinforcement Learning via Broadened ExplorationJian Hu, Mingjie Liu, Ximing Lu, Fang Wu et al.ICML 2026 · 16 citations
- Temperature as a Meta-Policy: Adaptive Temperature in LLM Reinforcement LearningHaoran Dang, Cuiling Lan, Hai Wan, Xibin Zhao et al.ICLR 2026 · 9 citations
- Optimal Self-Consistency for Efficient Reasoning with Large Language ModelsAustin Feng, Marius Alonso, Ambroise Odonnat, Vasilii Feofanov et al.ICML 2026 · 6 citations
- Large Language Models Develop Novel Social Biases Through Adaptive ExplorationAddison J. Wu, Ryan Liu, Xuechunzi Bai, Thomas GriffithsICML 2026 · 4 citations
- Demystifying LLM-as-a-Judge: Analytically Tractable Model for Inference-Time ScalingIndranil Halder, Cengiz PehlevanICML 2026 · 4 citations
Builds on12
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
Related papers
- Thermometer of Thoughts: Enhancing LLM's Exploration via Attention Temperature ModulationZhiyuan Yu, Shijian Xiao, Cam-Tu Nguyen, Zhangyue Yin et al.ACL 2026
- Future-Gain Guided Test-Time Learning for Large Language ModelsLangYu Bian, Jinwu Hu, Zitian Zhang, Dongjin Yang et al.ICML 2026 · 12 citations
- PRompt Optimization in Multi-Step Tasks (PROMST): Integrating Human Feedback and Heuristic-based SamplingYongchao Chen, Jacob Arkin, Yilun Hao, Yang Zhang et al.EMNLP 2024 · 6 citations
- In Good GRACES: Principled Teacher Selection for Knowledge DistillationAbhishek Panigrahi, Bingbin Liu, Sadhika Malladi, Sham M. Kakade et al.ICLR 2026 · 5 citations
- Balancing Act: Diversity and Consistency in Large Language Model EnsemblesAhmed Abdulaal, Chen Jin, Nina Montaña Brown, Aryo Pradipta Gema et al.ICLR 2025
