Unifying Attention Heads and Task Vectors via Hidden State Geometry in In-Context Learning
Haolin Yang, Hakaze Cho, Yiqiao Zhong, Naoya Inoue
Abstract
The unusual properties of in-context learning (ICL) have prompted investigations into the internal mechanisms of large language models. Prior work typically focuses on either special attention heads or task vectors at specific layers, but lacks a unified framework linking these components to the evolution of hidden states across layers that ultimately produce the model's output. In this paper, we propose such a framework for ICL primarily in classification tasks by analyzing two geometric factors that govern performance: the separability and alignment of query hidden states. A fine-grained analysis of layer-wise dynamics reveals a striking two-stage mechanism-separability emerges in early layers, while alignment develops in later layers. Ablation studies further show that Previous Token Heads drive separability, while Induction Heads and task vectors enhance alignment. Our findings thus bridge the gap between attention heads and task vectors, offering a unified account of ICL's underlying mechanisms. 1 is not a continuous function, S * as a supremum can be attained on S d-1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6e5657ef-52c8-48b2-8e50-815d1d4cde99Cited by top-tier papers5
- CoT Vectors: Transferring and Probing the Reasoning Mechanisms of LLMsLi Li, Ziyi Wang, Yongliang Wu, Jianfei Cai et al.ICLR 2026 · 4 citations
- Mechanism of Task-oriented Information Removal in In-context LearningHakaze Cho, Haolin Yang, Gouki Minegishi, Naoya InoueICLR 2026 · 3 citations
- Localizing Task Recognition and Task Learning in In-Context Learning via Attention Head AnalysisHaolin Yang, Hakaze Cho, Naoya InoueICLR 2026 · 2 citations
- HiFICL: High-Fidelity In-Context Learning for Multimodal TasksXiaoyu Li, Yuhang Liu, xuanshuo kang, zheng luo et al.CVPR 2026 · 1 citation
- Task Vectors, Learned Not Extracted: Performance Gains and Mechanistic InsightsHaolin Yang, Hakaze Cho, Kaize Ding, Naoya InoueICLR 2026
Builds on18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- A Survey on In-context LearningQingxiu Dong, Lei Li, Damai Dai, Ce Zheng et al.EMNLP 2024 · 479 citations
- Data Distributional Properties Drive Emergent In-Context Learning in TransformersStephanie C. Y. Chan, Adam Santoro, Andrew K. Lampinen, Jane X. Wang et al.NeurIPS 2022 · 407 citations
- Language Models Represent Space and TimeWes Gurnee, Max TegmarkICLR 2024 · 303 citations
- In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space SteeringSheng Liu, Haotian Ye, Lei Xing, James Y. ZouICML 2024 · 244 citations
Related papers
- Which Attention Heads Matter for In-Context Learning?Kayo Yin, Jacob SteinhardtICML 2025
- Unveiling Induction Heads: Provable Training Dynamics and Feature Learning in TransformersSiyu Chen, Heejune Sheen, Tianhao Wang, Zhuoran YangNeurIPS 2024 · 48 citations
- Beyond Induction Heads: In-Context Meta Learning Induces Multi-Phase Circuit EmergenceGouki Minegishi, Hiroki Furuta, Shohei Taniguchi, Yusuke Iwasawa et al.ICML 2025
- How do Large Language Models Learn In-Context? Query and Key Matrices of In-Context Heads are Two Towers for Metric LearningZeping Yu, Sophia AnaniadouEMNLP 2024 · 2 citations
- KV Shifting Attention Enhances Language ModelingMingyu Xu, Bingning Wang, Weipeng ChenICML 2025
