A Solvable High-Dimensional Model Where Nonlinear Autoencoders Learn Structure Invisible to PCA While Test Loss Misaligns With Generalization
Vicente Mendes, Lorenzo Bardone, Cédric Koller, Jorge Medina Moreira, Vittorio Erba, Emanuele Troiani, Lenka Zdeborova
Abstract
Many real-world datasets contain hidden structure that cannot be detected by simple linear correlations between input features. For example, latent factors may influence the data in a coordinated way, even though their effect is invisible to covariance-based methods such as PCA. In practice, nonlinear neural networks often succeed in extracting such hidden structure in unsupervised and self-supervised learning. However, constructing a minimal high-dimensional model where this advantage can be rigorously analyzed has remained an open theoretical challenge. We introduce a tractable high-dimensional spiked model with two latent factors: one visible to covariance, and one statistically dependent yet uncorrelated, appearing only in higher-order moments. PCA and linear autoencoders fail to recover the latter, while a minimal nonlinear autoencoder provably extracts both. We analyze both the population risk, and empirical risk minimization. Our model also provides a tractable example where self-supervised test loss is poorly aligned with representation quality: nonlinear autoencoders recover latent structure that linear methods miss, even though their reconstruction loss is higher.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7664451f-32f3-48ea-8e3d-e6d8e0212f4cCited by top-tier papers3
- Biased Generalization in Diffusion ModelsLuca Saglietti, Luca Biggio, Jerome Garnier-Brun, Davide Beltrame et al.ICML 2026 · 2 citations
- A theory of learning data statistics in diffusion models, from easy to hardLorenzo Bardone, Claudia Merger, Sebastian GoldtICML 2026
- A Fourier perspective on the learning dynamics of neural networks: from sample complexities to mechanistic insightsFabiola Ricci, Claudia Merger, Sebastian GoldtICML 2026
Builds on10
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- High-dimensional limit theorems for SGD: Effective dynamics and critical scalingGérard Ben Arous, Reza Gheissari, Aukosh JagannathNeurIPS 2022 · 94 citations
- Phase diagram of Stochastic Gradient Descent in high-dimensional two-layer neural networksRodrigo Veiga, Ludovic Stephan, Bruno Loureiro, Florent Krzakala et al.NeurIPS 2022 · 59 citations
- High-dimensional Asymptotics of Denoising AutoencodersHugo Cui, Lenka ZdeborováNeurIPS 2023 · 26 citations
- The dynamics of representation learning in shallow, non-linear autoencodersMaria Refinetti, Sebastian GoldtICML 2022 · 25 citations
Related papers
- Fundamental Limits of Two-layer Autoencoders, and Achieving Them with Gradient MethodsAleksandr Shevchenko, Kevin Kögler, Hamed Hassani, Marco MondelliICML 2023 · 3 citations
- A Random Matrix Theory of Masked Self-Supervised LearningArie Zurich, Federica Gerace, Bruno Loureiro, Yue LuICML 2026
- Understanding Masked Autoencoders via Hierarchical Latent Variable ModelsLingjing Kong, Martin Q. Ma, Guangyi Chen, Eric P. Xing et al.CVPR 2023
- Eliminating the Invariance on the Loss Landscape of Linear AutoencodersReza Oftadeh, Jiayi Shen, Zhangyang Wang, Dylan A. ShellICML 2020 · 12 citations
- Structure by Architecture: Structured Representations without RegularizationFelix Leeb, Giulia Lanzillotta, Yashas Annadani, Michel Besserve et al.ICLR 2023 · 1 citation
