Unveiling In-Context Learning: A Coordinate System to Understand Its Working Mechanism
Anhao Zhao, Fanghua Ye, Jinlan Fu, Xiaoyu Shen
摘要
Large language models (LLMs) exhibit remarkable in-context learning (ICL) capabilities. However, the underlying working mechanism of ICL remains poorly understood. Recent research presents two conflicting views on ICL: One emphasizes the impact of similar examples in the demonstrations, stressing the need for label correctness and more shots. The other attributes it to LLMs' inherent ability of task recognition, deeming label correctness and shot numbers of demonstrations as not crucial. In this work, we provide a Two-Dimensional Coordinate System that unifies both views into a systematic framework. The framework explains the behavior of ICL through two orthogonal variables: whether similar examples are presented in the demonstrations (perception) and whether LLMs can recognize the task (cognition). We propose the peak inverse rank metric to detect the task recognition ability of LLMs and study LLMs' reactions to different definitions of similarity. Based on these, we conduct extensive experiments to elucidate how ICL functions across each quadrant on multiple representative classification tasks. Finally, we extend our analyses to generation tasks, showing that our coordinate system can also be used to interpret ICL for generation tasks effectively. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Unraveling the Mechanics of Learning-Based Demonstration Selection for In-Context LearningHui Liu, Wenya Wang, Hao Sun, Chris Xing Tian 等ACL 2025 · 被引用 13 次
- Beyond Output Matching: Bidirectional Alignment for Enhanced In-Context LearningChengwei Qin, Wenhan Xia, Fangkai Jiao, Chen Chen 等ACL 2025 · 被引用 7 次
- TACO: Enhancing Multimodal In-context Learning via Task Mapping-Guided Sequence ConfigurationYanshu Li, Jianjiang Yang, Tian Yun, Pinyuan Feng 等EMNLP 2025 · 被引用 2 次
- Beyond Examples: Towards Automated Thought-level In-Context Reasoning for Large Language ModelsJinyang Wu, Mingkuan Feng, Shuai Zhang, Feihu Che 等ACL 2026
- Bridging Behavior and Semantics for Time-aware Cross-Domain Sequential RecommendationZhida Qin, Zemu Liu, Haoyan Fu, Chong Zhang 等SIGIR 2026
它引用的顶会 Paper21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Calibrate Before Use: Improving Few-shot Performance of Language ModelsZihao Zhao, Eric Wallace, Shi Feng, Dan Klein 等ICML 2021 · 被引用 1,843 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order SensitivityYao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel 等ACL 2022 · 被引用 1,494 次
相关 Paper
- What Do Language Models Learn in Context? The Structured Task HypothesisJiaoda Li, Yifan Hou, Mrinmaya Sachan, Ryan CotterellACL 2024 · 被引用 5 次
- Task Descriptors Help Transformers Learn Linear Models In-ContextRuomin Huang, Rong GeICLR 2025
- On Many-Shot In-Context Learning for Long-Context EvaluationKaijian Zou, Muhammad Khalifa, Lu WangACL 2025 · 被引用 6 次
- In-Context Learning Learns Label Relationships but Is Not Conventional LearningJannik Kossen, Yarin Gal, Tom RainforthICLR 2024 · 被引用 61 次
- Revisiting Demonstration Selection Strategies in In-Context LearningKeqin Peng, Liang Ding, Yancheng Yuan, Xuebo Liu 等ACL 2024
