To Cool or not to Cool? Temperature Network Meets Large Foundation Models via DRO
Zi-Hao Qiu, Siqi Guo, Mao Xu, Tuo Zhao, Lijun Zhang, Tianbao Yang
摘要
The temperature parameter plays a profound role during training and/or inference with large foundation models (LFMs) such as large language models (LLMs) and CLIP models. Particularly, it adjusts the logits in the softmax function in LLMs, which is crucial for next token generation, and it scales the similarities in the contrastive loss for training CLIP models. A significant question remains: "Is it viable to learn a neural network to predict a personalized temperature of any input data for enhancing LFMs?" In this paper, we present a principled framework for learning a small yet generalizable temperature prediction network (TempNet) to improve LFMs. Our solution is composed of a novel learning framework with a robust loss underpinned by constrained distributionally robust optimization (DRO), and a properly designed TempNet with theoretical inspiration. TempNet can be trained together with a large foundation model from scratch or learned separately given a pretrained foundation model. It is not only useful for predicting personalized temperature to promote the training of LFMs but also generalizable and transferable to new tasks. Our experiments on LLMs and CLIP models demonstrate that TempNet greatly improves the performance of existing solutions or models, e.g. Table 1 . The code to reproduce the experimental results in this paper can be found at https://github.com/zhqiu/TempNet .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Learning to Scale Logits for Temperature-Conditional GFlowNetsMinsu Kim, Joohwan Ko, Taeyoung Yun, Dinghuai Zhang 等ICML 2024 · 被引用 31 次
- Boosting Open Set Recognition Performance through Modulated Representation LearningAmit Kumar Kundu, Vaishnavi S Patil, Joseph JaJaICLR 2026 · 被引用 2 次
- Logits are All We Need to Adapt Closed ModelsGaurush Hiranandani, Haolun Wu, Subhojyoti Mukherjee, Sanmi KoyejoICML 2025
它引用的顶会 Paper17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty 等NeurIPS 2021 · 被引用 2,985 次
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
相关 Paper
- Model Steering: Learning with a Reference Model Improves Generalization Bounds and Scaling LawsXiyuan Wei, Ming Lin, Fanjiang Ye, Fengguang Song 等ICML 2025
- Not All Semantics are Created Equal: Contrastive Self-supervised Learning with Automatic Temperature IndividualizationZi-Hao Qiu, Quanqi Hu, Zhuoning Yuan, Denny Zhou 等ICML 2023 · 被引用 29 次
- Optimal Attention Temperature Improves the Robustness of In-Context Learning under Distribution Shift in High DimensionsSamet Demir, Zafer DoganICML 2026 · 被引用 1 次
- Temperature as a Meta-Policy: Adaptive Temperature in LLM Reinforcement LearningHaoran Dang, Cuiling Lan, Hai Wan, Xibin Zhao 等ICLR 2026 · 被引用 9 次
- PEARL: Towards Permutation-Resilient LLMsLiang Chen, Li Shen, Yang Deng, Xiaoyan Zhao 等ICLR 2025
