Provable In-Context Vector Arithmetic via Retrieving Task Concepts
Dake Bu, Wei Huang, Andi Han, Atsushi Nitanda, Qingfu Zhang, Hau-San Wong, Taiji Suzuki
摘要
In-context learning (ICL) has garnered significant attention for its ability to grasp functions/tasks from demonstrations. Recent studies suggest the presence of a latent task/function vector in LLMs during ICL. Merullo et al. (2024) showed that LLMs leverage this vector alongside the residual stream for Word2Veclike vector arithmetic, solving factual-recall ICL tasks. Additionally, recent work empirically highlighted the key role of Question-Answer data in enhancing factual-recall capabilities. Despite these insights, a theoretical explanation remains elusive. To move one step forward, we propose a theoretical framework building on empirically grounded hierarchical concept modeling. We develop an optimization theory, showing how nonlinear residual transformers trained via gradient descent on cross-entropy loss perform factualrecall ICL tasks via vector arithmetic. We prove 0-1 loss convergence and show the strong generalization, including robustness to concept recombination and distribution shifts. These results elucidate the advantages of transformers over static embedding predecessors. Empirical simulations corroborate our theoretical insights.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Causality ≠ Invariance: Function and Concept Vectors in LLMsGustaw Opielka, Hannes Rosenbusch, Claire E. StevensonICLR 2026 · 被引用 7 次
- Mechanism of Task-oriented Information Removal in In-context LearningHakaze Cho, Haolin Yang, Gouki Minegishi, Naoya InoueICLR 2026 · 被引用 3 次
- Trained Mamba Emulates Online Gradient Descent in In-Context Linear RegressionJiarui Jiang, Wei Huang, Miao Zhang, Taiji Suzuki 等NeurIPS 2025 · 被引用 2 次
- Towards a Theoretical Understanding of In-context Learning: Stability and Non-I.I.D GeneralisationYingjie Wang, Yutian Zhou, Shi Fu, Yuzhu Chen 等ICLR 2026
- Task Vectors, Learned Not Extracted: Performance Gains and Mechanistic InsightsHaolin Yang, Hakaze Cho, Kaize Ding, Naoya InoueICLR 2026
它引用的顶会 Paper33
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
- What Can Transformers Learn In-Context? A Case Study of Simple Function ClassesShivam Garg, Dimitris Tsipras, Percy Liang, Gregory ValiantNeurIPS 2022 · 被引用 883 次
- Merging Models with Fisher-Weighted AveragingMichael Matena, Colin RaffelNeurIPS 2022 · 被引用 741 次
- Transformers as Statisticians: Provable In-Context Learning with In-Context Algorithm SelectionYu Bai, Fan Chen, Huan Wang, Caiming Xiong 等NeurIPS 2023 · 被引用 356 次
- Physics of Language Models: Part 3.1, Knowledge Storage and ExtractionZeyuan Allen-Zhu, Yuanzhi LiICML 2024 · 被引用 258 次
相关 Paper
- In-Context Learning with Representations: Contextual Generalization of Trained TransformersTong Yang, Yu Huang, Yingbin Liang, Yuejie ChiNeurIPS 2024 · 被引用 45 次
- Transformers Efficiently Perform In-Context Logistic Regression via Normalized Gradient DescentChenyang Zhang, Yuan CaoICML 2026 · 被引用 1 次
- Function Vectors in Large Language ModelsEric Todd, Millicent L. Li, Arnab Sen Sharma, Aaron Mueller 等ICLR 2024 · 被引用 229 次
- Understanding Task Vectors in In-Context Learning: Emergence, Functionality, and LimitationsYuxin Dong, Jiachen Jiang, Zhihui Zhu, Xia NingICLR 2026 · 被引用 9 次
- Exact Conversion of In-Context Learning to Model Weights in Linearized-Attention TransformersBrian K. Chen, Tianyang Hu, Hui Jin, Hwee Kuan Lee 等ICML 2024 · 被引用 6 次
