Fundamental Limits of Two-layer Autoencoders, and Achieving Them with Gradient Methods
Aleksandr Shevchenko, Kevin Kögler, Hamed Hassani, Marco Mondelli
摘要
Autoencoders are a popular model in many branches of machine learning and lossy data compression. However, their fundamental limits, the performance of gradient methods and the features learnt during optimization remain poorly understood, even in the two-layer setting. In fact, earlier work has considered either linear autoencoders or specific training regimes (leading to vanishing or diverging compression rates). Our paper addresses this gap by focusing on non-linear two-layer autoencoders trained in the challenging proportional regime in which the input dimension scales linearly with the size of the representation. Our results characterize the minimizers of the population risk, and show that such minimizers are achieved by gradient methods; their structure is also unveiled, thus leading to a concise description of the features obtained via training. For the special case of a sign activation function, our analysis establishes the fundamental limits for the lossy compression of Gaussian sources via (shallow) autoencoders. Finally, while the results are proved for Gaussian data, numerical simulations on standard datasets display the universality of the theoretical predictions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- High-dimensional Asymptotics of Denoising AutoencodersHugo Cui, Lenka ZdeborováNeurIPS 2023 · 被引用 26 次
- Approaching Rate-Distortion Limits in Neural Compression with Lattice Transform CodingEric Lei, Hamed Hassani, Shirin Saeedi BidokhtiICLR 2025
- On the Feature Learning in Diffusion ModelsAndi Han, Wei Huang, Yuan Cao, Difan ZouICLR 2025
它引用的顶会 Paper7
- Learning curves of generic features maps for realistic datasets with a teacher-student modelBruno Loureiro, Cédric Gerbelot, Hugo Cui, Sebastian Goldt 等NeurIPS 2021 · 被引用 170 次
- Towards Empirical Sandwich Bounds on the Rate-Distortion FunctionYibo Yang, Stephan MandtICLR 2022 · 被引用 28 次
- The dynamics of representation learning in shallow, non-linear autoencodersMaria Refinetti, Sebastian GoldtICML 2022 · 被引用 25 次
- Eliminating the Invariance on the Loss Landscape of Linear AutoencodersReza Oftadeh, Jiayi Shen, Zhangyang Wang, Dylan A. ShellICML 2020 · 被引用 12 次
- Binary Iterative Hard Thresholding Converges with Optimal Number of Measurements for 1-Bit Compressed SensingNamiko Matsumoto, Arya MazumdarFOCS 2022 · 被引用 11 次
相关 Paper
- Compression of Structured Data with Autoencoders: Provable Benefit of Nonlinearities and DepthKevin Kögler, Aleksandr Shevchenko, Hamed Hassani, Marco MondelliICML 2024 · 被引用 2 次
- Implicit Rank-Minimizing AutoencoderLi Jing, Jure Zbontar, Yann LeCunNeurIPS 2020 · 被引用 63 次
- Implicit Bias of Large Depth Networks: a Notion of Rank for Nonlinear FunctionsArthur JacotICLR 2023 · 被引用 2 次
- The Usual Suspects? Reassessing Blame for VAE Posterior CollapseBin Dai, Ziyu Wang, David P. WipfICML 2020 · 被引用 89 次
- A Solvable High-Dimensional Model Where Nonlinear Autoencoders Learn Structure Invisible to PCA While Test Loss Misaligns With GeneralizationVicente Mendes, Lorenzo Bardone, Cédric Koller, Jorge Medina Moreira 等ICML 2026 · 被引用 6 次
