The Missing Alignment Link of In-context Learning on Sequences
Harshvardhan Agarwal, Sunita Sarawagi
Abstract
Large language models (LLMs) have demonstrated the capability to perform in-context learning (ICL) for completely unseen tasks in classification or language completion. Sequence to sequence (Seq2Seq) is another popular task category with several applications seeking quick adaptation with ICL. We present a systematic analysis of the ICL capability of LLMs on Seq2Seq tasks using a formal structured language-pair. Our study reveals a critical limitation: except for very short input sequences, ICL fails to achieve consistent learning across all output positions. This exposes a fundamental weakness of modern LLMs -their inability to effectively uncover the alignment between input and output sequences. Consequently, this limitation results in incomplete induction heads, which form the basis for in-context learning of new discrete mappings. To address these limitations, we propose ICA-Tune, a method for focused fine-tuning of an LLM using in-context examples. We present a mechanistic evaluation with two accuracy probes to show how alignment emerges in middle layers of an LLM without any direct supervision. This alignment leads to an abrupt jump in the completeness of the induction heads in higher layers. We show that compared to standard fine-tuning, ICA-Tune enables more sample efficient learning and generalizes better to OOD instances.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 919dbbd4-878e-483d-9ae9-88fa9b3b4babBuilds on22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- What Can Transformers Learn In-Context? A Case Study of Simple Function ClassesShivam Garg, Dimitris Tsipras, Percy Liang, Gregory ValiantNeurIPS 2022 · 883 citations
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento et al.ICML 2023 · 729 citations
- Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe et al.EMNLP 2022 · 634 citations
- Transformers learn to implement preconditioned gradient descent for in-context learningKwangjun Ahn, Xiang Cheng, Hadi Daneshmand, Suvrit SraNeurIPS 2023 · 324 citations
Related papers
- Meta-learning via Language Model In-context TuningYanda Chen, Ruiqi Zhong, Sheng Zha, George Karypis et al.ACL 2022
- What Do Language Models Learn in Context? The Structured Task HypothesisJiaoda Li, Yifan Hou, Mrinmaya Sachan, Ryan CotterellACL 2024 · 5 citations
- Fine-tuning vs. In-context Learning in Large Language Models: A Formal Language Learning PerspectiveBishwamittra Ghosh, Soumi Das, Till Speicher, Qinyuan Wu et al.ACL 2026
- Active Example Selection for In-Context LearningYiming Zhang, Shi Feng, Chenhao TanEMNLP 2022 · 84 citations
- Interpret and Improve In-Context Learning via the Lens of Input-Label MappingsChenghao Sun, Zhen Huang, Yonggang Zhang, Le Lu et al.ACL 2025 · 1 citation
