Unlabeled Data Can Provably Enhance In-Context Learning of Transformers
Renpu Liu, Jing Yang
Abstract
Large language models (LLMs) exhibit impressive in-context learning (ICL) capabilities, yet the quality of their predictions is fundamentally limited by the few costly labeled demonstrations that can fit into a prompt. Meanwhile, there exist vast and continuously growing amounts of unlabeled data that may be closely related to the ICL task. How to utilize such unlabeled data to provably enhance the performance of ICL thus becomes an emerging fundamental question. In this work, we propose a novel augmented ICL framework, in which the prompt includes a small set of labeled examples alongside a block of unlabeled inputs. We focus on the multi-class linear classification setting and demonstrate that, with chain-of-thought (CoT) prompting, a multi-layer transformer can effectively emulate an expectation-maximization (EM) algorithm. This enables the transformer to implicitly extract useful information from both labeled and unlabeled data, leading to provable improvements in ICL accuracy. Moreover, we show that such a transformer can be trained via teacher forcing, with its parameters converging to the desired solution at a linear rate. Experiments demonstrate that the augmented ICL framework consistently outperforms conventional few-shot ICL, providing empirical support for our theoretical findings. To the best of our knowledge, this is the first theoretical study on the impact of unlabeled data on the ICL performance of transformers.
2 Related Works ICL with Transformers. Brown et al. (2020) first shows that GPT-3, a transformer-based LLM, can perform new tasks from input-output pairs without parameter updates, suggesting its ICL ability. This intriguing phenomenon of transformers has attracted much attention, leading to various interpretations and hypotheses about its underlying mechanism. Research on ICL often demonstrates how transformers can emulate learning algorithms. For instance, several studies have designed transformers that execute gradient descent for linear and non-linear regression tasks (Akyürek et al., 2023;Von Oswald et al., 2023a). Recent works demonstrate that transformers can implement more advanced optimization algorithms other than vanilla gradient descent on various ICL tasks ( Baiet al., Recently, the training dynamics of transformers with CoT have been studied in Huang et al. (2025a) for weight prediction in linear regression, in Li et al. (2024a) for in-context supervised learning, in Kim and Suzuki (2025); Wen et al. (2025) for the parity problems, and in Huang et al. (2025b) for the even pairs problem. None of these studies, however, address whether the multi-step reasoning capacity through CoT can be utilized to extract information from unlabeled inputs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 36711063-9354-4d3d-81fe-b1e340bc6e9bCited by top-tier papers3
- When and How Unlabeled Data Provably Improve In-Context LearningYingcong Li, Xiangyu Chang, Muti Kara, Xiaofeng Liu et al.NeurIPS 2025 · 5 citations
- Learning to Evaluate: Cost-Effective Model Evaluation on Unlabeled Data with Meta-LearningTrinh Pham, Viet Huynh, Hongzhi Yin, Quoc Viet Hung Nguyen et al.KDD 2026 · 1 citation
- Symmetry Reveals the In-Context Classifier: Transformers Implement Mean-Shift DynamicsPatrick Lutz, Themistoklis Haris, Arjun Chandra, Aditya Gangrade et al.ICML 2026
Builds on40
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
Related papers
- Training Nonlinear Transformers for Chain-of-Thought Inference: A Theoretical Generalization AnalysisHongkang Li, Songtao Lu, Pin-Yu Chen, Xiaodong Cui et al.ICLR 2025
- To Think or Not to Think: The Hidden Cost of Meta-Training with Excessive CoT ExamplesVignesh Kothapalli, Ata Fatahi Baarzi, Hamed Firooz, Maziar SanjabiACL 2026
- Transformers Learn to Implement Multi-step Gradient Descent with Chain of ThoughtJianhao Huang, Zixuan Wang, Jason D. LeeICLR 2025
- Exact Conversion of In-Context Learning to Model Weights in Linearized-Attention TransformersBrian K. Chen, Tianyang Hu, Hui Jin, Hwee Kuan Lee et al.ICML 2024 · 6 citations
- Large Language Models Can Self-ImproveJiaxin Huang, Shixiang Gu, Le Hou, Yuexin Wu et al.EMNLP 2023 · 184 citations
