Transformers as Unsupervised Learning Algorithms: A study on Gaussian Mixtures
Zhiheng Chen, Ruofan Wu, Guanhua Fang
摘要
The transformer architecture has demonstrated remarkable capabilities in modern artificial intelligence, among which the capability of implicitly learning an internal model during inference time is widely believed to play a key role in the understanding of pre-trained large language models. However, most recent works have been focusing on studying supervised learning topics such as in-context learning, leaving the field of unsupervised learning largely unexplored. This paper investigates the capabilities of transformers in solving Gaussian Mixture Models (GMMs), a fundamental unsupervised learning problem through the lens of statistical estimation. We propose a transformer-based learning framework called Transformer for Gaussian Mixture Models (TGMM) that simultaneously learns to solve multiple GMM tasks using a shared transformer backbone. The learned models are empirically demonstrated to effectively mitigate the limitations of classical methods such as Expectation-Maximization (EM) or spectral algorithms, at the same time exhibit reasonable robustness to distribution shifts. Theoretically, we prove that transformers can efficiently approximate both the Expectation-Maximization (EM) algorithm and a core component of spectral methods—namely, cubic tensor power iterations. These results not only improve upon prior work on approximating the EM algorithm, but also provide, to our knowledge, the first theoretical guarantee that transformers can approximate high-order tensor operations. Our study bridges the gap between practical success and theoretical understanding, positioning transformers as versatile tools for unsupervised learning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space DualityTri Dao, Albert GuICML 2024 · 被引用 1,407 次
- An Explanation of In-context Learning as Implicit Bayesian InferenceSang Michael Xie, Aditi Raghunathan, Percy Liang, Tengyu MaICLR 2022 · 被引用 1,030 次
- What Can Transformers Learn In-Context? A Case Study of Simple Function ClassesShivam Garg, Dimitris Tsipras, Percy Liang, Gregory ValiantNeurIPS 2022 · 被引用 883 次
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento 等ICML 2023 · 被引用 729 次
相关 Paper
- Attention-based clusteringRodrigo Maulen-Soto, Pierre Marion, Claire BoyerNeurIPS 2025 · 被引用 3 次
- Distribution Transformers: Fast Approximate Bayesian Inference With On-The-Fly Prior AdaptationGeorge Whittle, Juliusz Ziomek, Jacob Rawling, Michael A OsborneICML 2026 · 被引用 13 次
- Transformers are almost optimal metalearners for linear classificationRoey Magen, Gal VardiNeurIPS 2025 · 被引用 2 次
- Unsupervised Meta-Learning via In-Context LearningAnna Vettoruzzo, Lorenzo Braccaioli, Joaquin Vanschoren, Marlena NowaczykICLR 2025
- A Transformer-based Framework for Multivariate Time Series Representation LearningGeorge Zerveas, Srideepika Jayaraman, Dhaval Patel, Anuradha Bhamidipaty 等KDD 2021 · 被引用 66 次
