When and How Unlabeled Data Provably Improve In-Context Learning
Yingcong Li, Xiangyu Chang, Muti Kara, Xiaofeng Liu, Amit K. Roy-Chowdhury, Samet Oymak
Abstract
Recent research shows that in-context learning (ICL) can be effective even when demonstrations have missing or incorrect labels. To shed light on this capability, we examine a canonical setting where the demonstrations are drawn according to a binary Gaussian mixture model (GMM) and a certain fraction of the demonstrations have missing labels. We provide a comprehensive theoretical study to show that: (1) The loss landscape of one-layer linear attention models recover the optimal fully-supervised estimator but completely fail to exploit unlabeled data; (2) In contrast, multilayer or looped transformers can effectively leverage unlabeled data by implicitly constructing estimators of the form with and denoting features and partially-observed labels (with missing entries set to zero). We characterize the class of polynomials that can be expressed as a function of depth and draw connections to Expectation Maximization, an iterative pseudo-labeling algorithm commonly used in semi-supervised learning. Importantly, the leading polynomial power is exponential in depth, so mild amount of depth/looping suffices. As an application of theory, we propose looping off-the-shelf tabular foundation models to enhance their semi-supervision capabilities. Extensive evaluations on real-world datasets show that our method significantly improves the semisupervised tabular learning performance over the standard single pass inference.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c9ab5c75-66ce-4a8e-9fa9-06b66ecd120bCited by top-tier papers6
- TabDPT: Scaling Tabular Foundation Models on Real DataJunwei Ma, Valentin Thomas, Rasa Hosseinzadeh, Alex Labach et al.NeurIPS 2025 · 118 citations
- Unlabeled Data Can Provably Enhance In-Context Learning of TransformersRenpu Liu, Jing YangNeurIPS 2025 · 3 citations
- Transformers are almost optimal metalearners for linear classificationRoey Magen, Gal VardiNeurIPS 2025 · 2 citations
- In Context Semi-Supervised LearningJiashuo Fan, Paul Rosu, Aaron T. Wang, Lawrence Carin et al.ICLR 2026 · 2 citations
- Symmetry Reveals the In-Context Classifier: Transformers Implement Mean-Shift DynamicsPatrick Lutz, Themistoklis Haris, Arjun Chandra, Aditya Gangrade et al.ICML 2026
Builds on24
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Explanation of In-context Learning as Implicit Bayesian InferenceSang Michael Xie, Aditi Raghunathan, Percy Liang, Tengyu MaICLR 2022 · 1,030 citations
- What Can Transformers Learn In-Context? A Case Study of Simple Function ClassesShivam Garg, Dimitris Tsipras, Percy Liang, Gregory ValiantNeurIPS 2022 · 883 citations
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento et al.ICML 2023 · 729 citations
- Transformers as Statisticians: Provable In-Context Learning with In-Context Algorithm SelectionYu Bai, Fan Chen, Huan Wang, Caiming Xiong et al.NeurIPS 2023 · 356 citations
Related papers
- Transformers Efficiently Perform In-Context Logistic Regression via Normalized Gradient DescentChenyang Zhang, Yuan CaoICML 2026 · 1 citation
- Is One Layer Enough? Understanding Inference Dynamics in Tabular Foundation ModelsAmir Rezaei Balef, Mykhailo Koshil, Katharina EggenspergerICML 2026 · 2 citations
- In-Context Learning with Representations: Contextual Generalization of Trained TransformersTong Yang, Yu Huang, Yingbin Liang, Yuejie ChiNeurIPS 2024 · 45 citations
- On the Training Convergence of Transformers for In-Context Classification of Gaussian MixturesWei Shen, Ruida Zhou, Jing Yang, Cong ShenICML 2025
- TabICL: A Tabular Foundation Model for In-Context Learning on Large DataJingang Qu, David Holzmüller, Gaël Varoquaux, Marine Le MorvanICML 2025
