Revisiting In-context Learning Inference Circuit in Large Language Models
Hakaze Cho, Mariko Kato, Yoshihiro Sakai, Naoya Inoue
Abstract
In-context Learning (ICL) is an emerging few-shot learning paradigm on Language Models (LMs) with inner mechanisms un-explored. There are already existing works describing the inner processing of ICL, while they struggle to capture all the inference phenomena in large language models. Therefore, this paper proposes a comprehensive circuit to model the inference dynamics and try to explain the observed phenomena of ICL. In detail, we divide ICL inference into 3 major operations: (1) Input Text Encode: LMs encode every input text (in the demonstrations and queries) into linear representation in the hidden states with sufficient information to solve ICL tasks. (2) Semantics Merge: LMs merge the encoded representations of demonstrations with their corresponding label tokens to produce joint representations of labels and demonstrations. (3) Feature Retrieval and Copy: LMs search the joint representations of demonstrations similar to the query representation on a task subspace, and copy the searched representations into the query. Then, language model heads capture these copied label representations to a certain extent and decode them into predicted labels. Through careful measurements, the proposed inference circuit successfully captures and unifies many fragmented phenomena observed during the ICL process, making it a comprehensive and practical explanation of the ICL inference process. Moreover, ablation analysis by disabling the proposed steps seriously damages the ICL performance, suggesting the proposed inference circuit is a dominating mechanism. Additionally, we confirm and list some bypass mechanisms that solve ICL tasks in parallel with the proposed circuit.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6e4fcb9f-7471-4935-b85a-b892d276cb71Cited by top-tier papers9
- Unifying Attention Heads and Task Vectors via Hidden State Geometry in In-Context LearningHaolin Yang, Hakaze Cho, Yiqiao Zhong, Naoya InoueNeurIPS 2025 · 11 citations
- SafeSeek: Universal Attribution of Safety Circuits in Language ModelsMiao Yu, Siyuan Fu, Moayad Aloqaily, Zhenhong Zhou et al.ICML 2026 · 3 citations
- Mechanism of Task-oriented Information Removal in In-context LearningHakaze Cho, Haolin Yang, Gouki Minegishi, Naoya InoueICLR 2026 · 3 citations
- Localizing Task Recognition and Task Learning in In-Context Learning via Attention Head AnalysisHaolin Yang, Hakaze Cho, Naoya InoueICLR 2026 · 2 citations
- How Few-Shot Examples Add Up: A Causal Decomposition of Function Vectors in In-Context LearningEntang Wang, Yiwei Wang, Aleksandra Bakalova, Michael HahnICML 2026 · 1 citation
Related papers
- What Do Language Models Learn in Context? The Structured Task HypothesisJiaoda Li, Yifan Hou, Mrinmaya Sachan, Ryan CotterellACL 2024 · 5 citations
- Why Multimodal In-Context Learning Lags Behind? Unveiling the Inner Mechanisms and BottlenecksYu Wang, Sharon LiACL 2026
- Competition Dynamics Shape Algorithmic Phases of In-Context LearningCore Francisco Park, Ekdeep Singh Lubana, Hidenori TanakaICLR 2025
- Which Attention Heads Matter for In-Context Learning?Kayo Yin, Jacob SteinhardtICML 2025
- Unraveling the Mechanics of Learning-Based Demonstration Selection for In-Context LearningHui Liu, Wenya Wang, Hao Sun, Chris Xing Tian et al.ACL 2025 · 13 citations
