To Cool or not to Cool? Temperature Network Meets Large Foundation Models via DRO
Zi-Hao Qiu, Siqi Guo, Mao Xu, Tuo Zhao, Lijun Zhang, Tianbao Yang
Abstract
The temperature parameter plays a profound role during training and/or inference with large foundation models (LFMs) such as large language models (LLMs) and CLIP models. Particularly, it adjusts the logits in the softmax function in LLMs, which is crucial for next token generation, and it scales the similarities in the contrastive loss for training CLIP models. A significant question remains: "Is it viable to learn a neural network to predict a personalized temperature of any input data for enhancing LFMs?" In this paper, we present a principled framework for learning a small yet generalizable temperature prediction network (TempNet) to improve LFMs. Our solution is composed of a novel learning framework with a robust loss underpinned by constrained distributionally robust optimization (DRO), and a properly designed TempNet with theoretical inspiration. TempNet can be trained together with a large foundation model from scratch or learned separately given a pretrained foundation model. It is not only useful for predicting personalized temperature to promote the training of LFMs but also generalizable and transferable to new tasks. Our experiments on LLMs and CLIP models demonstrate that TempNet greatly improves the performance of existing solutions or models, e.g. Table 1 . The code to reproduce the experimental results in this paper can be found at https://github.com/zhqiu/TempNet .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eca690bf-0224-44b1-bc8b-8b1f7e55d926Cited by top-tier papers3
- Learning to Scale Logits for Temperature-Conditional GFlowNetsMinsu Kim, Joohwan Ko, Taeyoung Yun, Dinghuai Zhang et al.ICML 2024 · 31 citations
- Boosting Open Set Recognition Performance through Modulated Representation LearningAmit Kumar Kundu, Vaishnavi S Patil, Joseph JaJaICLR 2026 · 2 citations
- Logits are All We Need to Adapt Closed ModelsGaurush Hiranandani, Haolun Wu, Subhojyoti Mukherjee, Sanmi KoyejoICML 2025
Builds on17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
Related papers
- Model Steering: Learning with a Reference Model Improves Generalization Bounds and Scaling LawsXiyuan Wei, Ming Lin, Fanjiang Ye, Fengguang Song et al.ICML 2025
- Not All Semantics are Created Equal: Contrastive Self-supervised Learning with Automatic Temperature IndividualizationZi-Hao Qiu, Quanqi Hu, Zhuoning Yuan, Denny Zhou et al.ICML 2023 · 29 citations
- Optimal Attention Temperature Improves the Robustness of In-Context Learning under Distribution Shift in High DimensionsSamet Demir, Zafer DoganICML 2026 · 1 citation
- Temperature as a Meta-Policy: Adaptive Temperature in LLM Reinforcement LearningHaoran Dang, Cuiling Lan, Hai Wan, Xibin Zhao et al.ICLR 2026 · 9 citations
- PEARL: Towards Permutation-Resilient LLMsLiang Chen, Li Shen, Yang Deng, Xiaoyan Zhao et al.ICLR 2025
