Optimal Attention Temperature Improves the Robustness of In-Context Learning under Distribution Shift in High Dimensions
Samet Demir, Zafer Dogan
Abstract
Pretrained Transformers can perform in-context learning (ICL) from a few demonstrations, but this ability can fail sharply when the test distribution differs from pretraining-a common deployment setting. We study attention temperature as a simple inference-time control for improving ICL robustness under such shifts. In a highdimensional linear-regression framework, we analyze a Transformer with "approximate softmax" attention, which preserves softmax's normalization and temperature-dependent selectivity while remaining tractable. We derive a closed-form expression for the ICL generalization error under distribution shift, and show that it is minimized by an explicit optimal attention temperature. This characterization yields interpretable guidance by linking the best temperature to moments of the pre-softmax attention scores, and predicts when temperature adjustment can recover near Bayesoptimal performance. We validate the theory with extensive simulations, and further demonstrate gains on pretrained LLMs (GPT-2 and Llama2-7B) on question-answering benchmarks under distribution shift induced by noisy in-context demonstrations. Overall, attention temperature emerges as a principled, lightweight knob for improving the robustness of ICL in pretrained Transformers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dcced913-ce25-4c5c-8ef9-cf9b967abb37Cited by top-tier papers1
Ask how each one uses itBuilds on21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- What Can Transformers Learn In-Context? A Case Study of Simple Function ClassesShivam Garg, Dimitris Tsipras, Percy Liang, Gregory ValiantNeurIPS 2022 · 883 citations
- Are Emergent Abilities of Large Language Models a Mirage?Rylan Schaeffer, Brando Miranda, Sanmi KoyejoNeurIPS 2023 · 796 citations
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento et al.ICML 2023 · 729 citations
- Transformers as Statisticians: Provable In-Context Learning with In-Context Algorithm SelectionYu Bai, Fan Chen, Huan Wang, Caiming Xiong et al.NeurIPS 2023 · 356 citations
Related papers
- SSA: Improving Performance With a Better Scoring FunctionOmar Naim, Swarnadeep Bhar, Jérôme Bolte, Nicholas AsherACL 2026
- Towards Understanding How Transformers Learn In-context Through a Representation Learning LensRuifeng Ren, Yong LiuNeurIPS 2024 · 26 citations
- How Does the Pretraining Distribution Shape In-Context Learning? A Fundamental Trade-OffWaïss Azizian, Ali HasanICML 2026
- In-Context Learning with Transformers: Softmax Attention Adapts to Function LipschitznessLiam Collins, Advait Parulekar, Aryan Mokhtari, Sujay Sanghavi et al.NeurIPS 2024 · 33 citations
- When Parts Are Greater Than Sums: Individual LLM Components Can Outperform Full ModelsTing-Yun Chang, Jesse Thomason, Robin JiaEMNLP 2024
