High-dimensional Asymptotics of Denoising Autoencoders
Hugo Cui, Lenka Zdeborová
Abstract
We address the problem of denoising data from a Gaussian mixture using a two-layer non-linear autoencoder with tied weights and a skip connection. We consider the high-dimensional limit where the number of training samples and the input dimension jointly tend to infinity while the number of hidden units remains bounded. We provide closed-form expressions for the denoising mean-squared test error. Building on this result, we quantitatively characterize the advantage of the considered architecture over the autoencoder without the skip connection that relates closely to principal component analysis. We further show that our results accurately capture the learning curves on a range of real data sets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cd13e18f-94a1-48f9-a1a9-5d5440bbb8fcCited by top-tier papers12
- A Phase Transition between Positional and Semantic Learning in a Solvable Model of Dot-Product AttentionHugo Cui, Freya Behrens, Florent Krzakala, Lenka ZdeborováNeurIPS 2024 · 35 citations
- Analysis of Learning a Flow-based Generative Model from Limited Sample ComplexityHugo Cui, Florent Krzakala, Eric Vanden-Eijnden, Lenka ZdeborováICLR 2024 · 31 citations
- Asymptotics of feature learning in two-layer networks after one gradient-stepHugo Cui, Luca Pesce, Yatin Dandi, Florent Krzakala et al.ICML 2024 · 30 citations
- Generalization of Diffusion Models Arises with a Balanced Representation SpaceZekai Zhang, Xiao Li, Xiang Li, Lianghe Shi et al.ICLR 2026 · 14 citations
- A solvable model of learning generative diffusion: theory and insightsHugo Cui, Cengiz Pehlevan, Yue M. LuNeurIPS 2025 · 11 citations
Builds on9
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Spectrum Dependent Learning Curves in Kernel Regression and Wide Neural NetworksBlake Bordelon, Abdulkadir Canatar, Cengiz PehlevanICML 2020 · 245 citations
- Generalisation error in learning with random features and the hidden manifold modelFederica Gerace, Bruno Loureiro, Florent Krzakala, Marc Mézard et al.ICML 2020 · 184 citations
- Learning curves of generic features maps for realistic datasets with a teacher-student modelBruno Loureiro, Cédric Gerbelot, Hugo Cui, Sebastian Goldt et al.NeurIPS 2021 · 170 citations
- Generalization error in high-dimensional perceptrons: Approaching Bayes error with convex optimizationBenjamin Aubin, Florent Krzakala, Yue M. Lu, Lenka ZdeborováNeurIPS 2020 · 67 citations
Related papers
- The dynamics of representation learning in shallow, non-linear autoencodersMaria Refinetti, Sebastian GoldtICML 2022 · 25 citations
- A Solvable High-Dimensional Model Where Nonlinear Autoencoders Learn Structure Invisible to PCA While Test Loss Misaligns With GeneralizationVicente Mendes, Lorenzo Bardone, Cédric Koller, Jorge Medina Moreira et al.ICML 2026 · 6 citations
- Fundamental Limits of Two-layer Autoencoders, and Achieving Them with Gradient MethodsAleksandr Shevchenko, Kevin Kögler, Hamed Hassani, Marco MondelliICML 2023 · 3 citations
- Autoencoders that don't overfit towards the IdentityHarald SteckNeurIPS 2020 · 72 citations
- Compression of Structured Data with Autoencoders: Provable Benefit of Nonlinearities and DepthKevin Kögler, Aleksandr Shevchenko, Hamed Hassani, Marco MondelliICML 2024 · 2 citations
