Learning with Restricted Boltzmann Machines: Asymptotics of AMP and GD in High Dimensions
Yizhou Xu, Florent Krzakala, Lenka Zdeborová
摘要
The Restricted Boltzmann Machine (RBM) is one of the simplest generative neural networks capable of learning input distributions. Despite its simplicity, the analysis of its performance in learning from the training data is only well understood in cases that essentially reduce to singular value decomposition of the data. Here, we consider the limit of a large dimension of the input space and a constant number of hidden units. In this limit, we simplify the standard RBM training objective into a form that is equivalent to the multi-index model with non-separable regularization. This opens a path to analyze training of the RBM using methods that are established for multi-index models, such as Approximate Message Passing (AMP) and its state evolution, and the analysis of Gradient Descent (GD) via the dynamical mean-field theory. We then give rigorous asymptotics of the training dynamics of RBMs on data generated by the spiked covariance model as a prototype of a structure suitable for unsupervised learning. We show in particular that RBMs reach the optimal computational weak recovery threshold, aligning with the Baik-Ben Arous-Péché (BBP) transition, in the spiked covariance model.
replica method also exist in the statistical physics literature [Alemanno et al., 2023, Thériault et al., 2024, Manzan and Tantari, 2024], but in addition to their non-rigorous nature, they often fail to reflect the actual training procedure used in practice-namely, maximum likelihood estimation. Instead, they focus on Bayesian posterior sampling, which yields qualitatively different predictions. The only works studying the full learning on the usual likelihood objective is as far as we know [Bachtis et al., 2024, Harsh et al., 2020], and we will comment on the differences with our work below. This gap in our theoretical understanding of RBMs stands in stark contrast to recent progress on feedforward networks, a field that has seen a surge of analytical results, especially in high-dimensional regimes with synthetic data, where rigorous insights are increasingly available, see e.g.,Soltanolkotabi et al. [2018], Mei et al. [2019], Celentano et al. [2020], Gerbelot et al. [2024], Bietti et al. [2023].
Motivated by this progress, we analyze empirical likelihood maximization for RBMs directly, in a setting that mirrors practical training procedures. Specifically, we consider the high-dimensional limit in which both the data dimension and the number of samples tend to infinity, while the number of hidden units remains fixed. This regime is relevant for many real-world applications, where data often lies on low-dimensional manifolds embedded in high-dimensional spaces. In this limit, the likelihood function admits a simplified form that reduces to an unsupervised variant of the multi-index model-a framework that has proven useful for studying neural networks and gradient descent in non-convex, high-dimensional settings [Saad and Solla, 1995, Abbe et al., 2023, Damian et al., 2023, Troiani et al., 2024]. Leveraging this connection, one can characterize the learning dynamics and derive sharp asymptotic results for data generated by the celebrated spiked covariance model [Johnstone, 2001].
Main results -Our contributions are the following: i) We prove that, in high dimensions and with finitely many hidden units, the RBM training objective is asymptotically equivalent to an unsupervised multi-index model with non-separable regularization. This characterization is key to our analysis, as it allows the use of established mathematical techniques for synthetic data. As a model for data, we consider the well-studied spiked covariance model and emphasise its equivalence to a teacher RBM. ii) Thanks to this mapping, we show that one can rigorously analyze the asymptotics of RBMs trained on the spiked covariance data. We apply a large body of recent results from multi-index and supervised models to RBMs. In particular, we derive sharp asymptotics for the simplified objective and introduce an Approximate Message Passing (AMP) algorithm to train the model. This allows us to characterize the global optimum of the objective for a number of cases. We establish that the weak recovery threshold of RBMs matches the optimal BBP threshold ([Baik et al., 2005], i.e., the minimal signal-to-noise ratio to detect the signal). iii) We provide rigorous, closed-form equations describing the exact high-dimensional asymptotics of gradient descent dynamics. These results generalize dynamical mean-field theory, which was previously limited to multi-index and tensor learning problems.
Overall, our analysis opens the door to a precise mathematical understanding of unsupervised learning in RBMs, a foundational generative model. By mapping the likelihood landscape of RBMs in the high-dimensional regime to that of an effective model, we are able to import a range of tools and results that were previously restricted to supervised learning in single-and two-layer networks. T
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Dynamical mean-field theory for stochastic gradient descent in Gaussian mixture classificationFrancesca Mignacco, Florent Krzakala, Pierfrancesco Urbani, Lenka ZdeborováNeurIPS 2020 · 被引用 95 次
- Smoothing the Landscape Boosts the Signal for SGD: Optimal Sample Complexity for Learning Single Index ModelsAlex Damian, Eshaan Nichani, Rong Ge, Jason D. LeeNeurIPS 2023 · 被引用 67 次
- Generalization Error of Generalized Linear Models in High DimensionsMelikasadat Emami, Mojtaba Sahraee-Ardakan, Parthe Pandit, Sundeep Rangan 等ICML 2020 · 被引用 40 次
- Cascade of phase transitions in the training of energy-based modelsDimitrios Bachtis, Giulio Biroli, Aurélien Decelle, Beatriz SeoaneNeurIPS 2024 · 被引用 17 次
相关 Paper
- Equilibrium and non-Equilibrium regimes in the learning of Restricted Boltzmann MachinesAurélien Decelle, Cyril Furtlehner, Beatriz SeoaneNeurIPS 2021 · 被引用 40 次
- Optimal Spectral Transitions in High-Dimensional Multi-Index ModelsLeonardo Defilippis, Yatin Dandi, Pierre Mergny, Florent Krzakala 等NeurIPS 2025 · 被引用 8 次
- From Boltzmann Machines to Neural Networks and Back AgainSurbhi Goel, Adam R. Klivans, Frederic KoehlerNeurIPS 2020 · 被引用 7 次
- Fast training and sampling of Restricted Boltzmann MachinesNicolas Béreux, Aurélien Decelle, Cyril Furtlehner, Lorenzo Rosset 等ICLR 2025 · 被引用 2 次
- A Random Matrix Theory of Masked Self-Supervised LearningArie Zurich, Federica Gerace, Bruno Loureiro, Yue LuICML 2026
