Optimizing Temperature for Language Models with Multi-Sample Inference
Weihua Du, Yiming Yang, Sean Welleck
摘要
Multi-sample aggregation strategies, such as majority voting and best-of-N sampling, are widely used in contemporary large language models (LLMs) to enhance predictive accuracy across various tasks. A key challenge in this process is temperature selection, which significantly impacts model performance. Existing approaches either rely on a fixed default temperature or require labeled validation data for tuning, which are often scarce and difficult to obtain. This paper addresses the challenge of automatically identifying the (near)-optimal temperature for different LLMs using multi-sample aggregation strategies, without relying on task-specific validation data. We provide a comprehensive analysis of temperature's role in performance optimization, considering variations in model architectures, datasets, task types, model sizes, and predictive accuracy. Furthermore, we propose a novel entropy-based metric for automated temperature optimization, which consistently outperforms fixed-temperature baselines. Additionally, we incorporate a stochastic process model to enhance interpretability, offering deeper insights into the relationship between temperature and model performance. Our code is available at https://github.com/StigLidu/dualdistill . Optimizing Temperature for Language Models with Multi-Sample Inference ing how to optimize the sampling process to enhance LLM performance under different conditions, including variations in training datasets, task types, and model sizes. A crucial open question is how to tune temperature, a key hyperparameter that controls the smoothness of the system-learned distribution. Intuitively, increasing the temperature leads to a smoother distribution, enhancing the diversity of sampled outputs. However, excessively high temperatures can introduce many low-quality samples, making aggregation more challenging (Holtzman et al., 2019; Renze & Guven, 2024) . Conversely, lowering the temperature results in a highly concentrated distribution, reducing diversity and potentially omitting high-quality samples. Striking the right balance between over-sampling and under-sampling is therefore essential for optimizing LLM performance. A common practice in prior evaluations is to use the same temperature across all methods despite variations in training datasets, task types, model sizes, and aggregation strategies. This practice is clearly suboptimal. An alternative approach is to empirically tune the temperature using labeled validation data for each task, dataset, model size, and aggregation strategy (Zhang et al., 2024a; Dhuliawala et al., 2024) . However, such a process is tedious and time-consuming and heavily dependent on the availability of labeled validation data, limiting its applicability when such data are scarce. In this paper, we present the first systematic investigation of how temperature affects LLM performance under multisample aggregation strategies across various conditions. Furthermore, we propose a principled algorithmic solution for automated temperature optimization without requiring labeled validation data. Our key idea is as follows: 1. We use the confidence score of each model as a selfassessment measure.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- BroRL: Scaling Reinforcement Learning via Broadened ExplorationJian Hu, Mingjie Liu, Ximing Lu, Fang Wu 等ICML 2026 · 被引用 16 次
- Temperature as a Meta-Policy: Adaptive Temperature in LLM Reinforcement LearningHaoran Dang, Cuiling Lan, Hai Wan, Xibin Zhao 等ICLR 2026 · 被引用 9 次
- Optimal Self-Consistency for Efficient Reasoning with Large Language ModelsAustin Feng, Marius Alonso, Ambroise Odonnat, Vasilii Feofanov 等ICML 2026 · 被引用 6 次
- Large Language Models Develop Novel Social Biases Through Adaptive ExplorationAddison J. Wu, Ryan Liu, Xuechunzi Bai, Thomas GriffithsICML 2026 · 被引用 4 次
- Demystifying LLM-as-a-Judge: Analytically Tractable Model for Inference-Time ScalingIndranil Halder, Cengiz PehlevanICML 2026 · 被引用 4 次
它引用的顶会 Paper12
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
相关 Paper
- Thermometer of Thoughts: Enhancing LLM's Exploration via Attention Temperature ModulationZhiyuan Yu, Shijian Xiao, Cam-Tu Nguyen, Zhangyue Yin 等ACL 2026
- Future-Gain Guided Test-Time Learning for Large Language ModelsLangYu Bian, Jinwu Hu, Zitian Zhang, Dongjin Yang 等ICML 2026 · 被引用 12 次
- PRompt Optimization in Multi-Step Tasks (PROMST): Integrating Human Feedback and Heuristic-based SamplingYongchao Chen, Jacob Arkin, Yilun Hao, Yang Zhang 等EMNLP 2024 · 被引用 6 次
- In Good GRACES: Principled Teacher Selection for Knowledge DistillationAbhishek Panigrahi, Bingbin Liu, Sadhika Malladi, Sham M. Kakade 等ICLR 2026 · 被引用 5 次
- Balancing Act: Diversity and Consistency in Large Language Model EnsemblesAhmed Abdulaal, Chen Jin, Nina Montaña Brown, Aryo Pradipta Gema 等ICLR 2025
