Where does In-context Learning Happen in Large Language Models?
Suzanna Sia, David Mueller, Kevin Duh
Abstract
Self-supervised large language models have demonstrated the ability to perform various tasks via in-context learning, but little is known about where the model locates the task with respect to prompt instructions and demonstration examples. In this work, we attempt to characterize the region where large language models transition from recognizing the task to performing the task. Through a series of layer-wise context-masking experiments on GPTN EO 2.7B, B LOOM 3B, and S TARCODER 2-7B, L LAMA 3.1-8B, L LAMA 3.1-8B-I NSTRUCT , on Machine Translation and Code generation, we demonstrate evidence of a "task recognition" point where the task is encoded into the input representations and attention to context is no longer necessary. Taking advantage of this redundancy results in 45% computational savings when prompting with 5 examples, and task recognition achieved at layer 14 / 32 using an example with Machine Translation. Our findings also have implication for resource and parameter efficient fine-tuning; we observe a correspondence between fine-tuning performance of individual LoRA layers and the task recognition layers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 54120624-246d-4414-941b-171eff0fd792Cited by top-tier papers3
- Unifying Attention Heads and Task Vectors via Hidden State Geometry in In-Context LearningHaolin Yang, Hakaze Cho, Yiqiao Zhong, Naoya InoueNeurIPS 2025 · 11 citations
- Internal Chain-of-Thought: Empirical Evidence for Layer-wise Subtask Scheduling in LLMsZhipeng Yang, Junzhuo Li, Siyu Xia, Xuming HuEMNLP 2025
- On the Role of Model Prior in Real-World Inductive ReasoningZhuo Liu, Ding Yu, Hangfeng HeEMNLP 2025
Builds on24
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach et al.ICLR 2022 · 1,976 citations
- Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order SensitivityYao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel et al.ACL 2022 · 1,494 citations
Related papers
- LongLoRA: Efficient Fine-tuning of Long-Context Large Language ModelsYukang Chen, Shengju Qian, Haotian Tang, Xin Lai et al.ICLR 2024 · 254 citations
- Emergence and Effectiveness of Task Vectors in In-Context Learning: An Encoder Decoder PerspectiveSeungwook Han, Jinyeop Song, Jeff Gore, Pulkit AgrawalICML 2025
- LoRA-Gen: Specializing Large Language Model via Online LoRA GenerationYicheng Xiao, Lin Song, Rui Yan, Cheng Cheng et al.ICML 2025
- Dodo: Dynamic Contextual Compression for Decoder-only LMsGuanghui Qin, Corby Rosset, Ethan C. Chau, Nikhil Rao et al.ACL 2024
- From Bottom to Top: Extending the Potential of Parameter Efficient Fine-TuningJihao Gu, Zelin Wang, Yibo Zhang, Ziji Zhang et al.EMNLP 2024 · 3 citations
