Are Gaussian Data All You Need? The Extents and Limits of Universality in High-Dimensional Generalized Linear Estimation
Luca Pesce, Florent Krzakala, Bruno Loureiro, Ludovic Stephan
摘要
In this manuscript we consider the problem of generalized linear estimation on Gaussian mixture data with labels given by a single-index model. Our first result is a sharp asymptotic expression for the test and training errors in the high-dimensional regime. Motivated by the recent stream of results on the Gaussian universality of the test and training errors in generalized linear estimation, we ask ourselves the question: "when is a single Gaussian enough to characterize the error?". Our formula allow us to give sharp answers to this question, both in the positive and negative directions. More precisely, we show that the sufficient conditions for Gaussian universality (or lack of thereof) crucially depend on the alignment between the target weights and the means and covariances of the mixture clusters, which we precisely quantify. In the particular case of least-squares interpolation, we prove a strong universality property of the training error, and show it follows a simple, closed-form expression. Finally, we apply our results to real datasets, clarifying some recent discussion in the literature about Gaussian universality of the errors in this context.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Asymptotics of feature learning in two-layer networks after one gradient-stepHugo Cui, Luca Pesce, Yatin Dandi, Florent Krzakala 等ICML 2024 · 被引用 30 次
- Restoring balance: principled under/oversampling of data for optimal classificationEmanuele Loffredo, Mauro Pastore, Simona Cocco, Rémi MonassonICML 2024 · 被引用 13 次
- Asymptotics of Learning with Deep Structured (Random) FeaturesDominik Schröder, Daniil Dmitriev, Hugo Cui, Bruno LoureiroICML 2024 · 被引用 12 次
- The Effect of Optimal Self-Distillation in Noisy Gaussian Mixture ModelKaito Takanami, Takashi Takahashi, Ayaka SakataNeurIPS 2025 · 被引用 4 次
- A Noise Sensitivity Exponent Controls Large Statistical-to-Computational Gaps in Single- and Multi-Index ModelsLeonardo Defilippis, FLORENT KRZAKALA, Bruno Loureiro, Antoine MaillardICML 2026 · 被引用 1 次
它引用的顶会 Paper10
- Spectrum Dependent Learning Curves in Kernel Regression and Wide Neural NetworksBlake Bordelon, Abdulkadir Canatar, Cengiz PehlevanICML 2020 · 被引用 245 次
- Generalisation error in learning with random features and the hidden manifold modelFederica Gerace, Bruno Loureiro, Florent Krzakala, Marc Mézard 等ICML 2020 · 被引用 184 次
- High-dimensional Asymptotics of Feature Learning: How One Gradient Step Improves the RepresentationJimmy Ba, Murat A. Erdogdu, Taiji Suzuki, Zhichao Wang 等NeurIPS 2022 · 被引用 173 次
- Learning curves of generic features maps for realistic datasets with a teacher-student modelBruno Loureiro, Cédric Gerbelot, Hugo Cui, Sebastian Goldt 等NeurIPS 2021 · 被引用 170 次
- On the Optimal Weighted Regularization in Overparameterized Linear RegressionDenny Wu, Ji XuNeurIPS 2020 · 被引用 151 次
相关 Paper
- The Breakdown of Gaussian Universality in Classification of High-dimensional Linear Factor MixturesXiaoyi Mai, Zhenyu LiaoICLR 2025
- Universality laws for Gaussian mixtures in generalized linear modelsYatin Dandi, Ludovic Stephan, Florent Krzakala, Bruno Loureiro 等NeurIPS 2023 · 被引用 40 次
- Asymptotic Analysis of Two-Layer Neural Networks after One Gradient Step under Gaussian Mixtures Data with StructureSamet Demir, Zafer DoganICLR 2025
- Robust group and simultaneous inferences for high-dimensional single index modelWeichao Yang, Hongwei Shi, Xu Guo, Changliang ZouNeurIPS 2024 · 被引用 2 次
- Learning Gaussian Mixtures with Generalized Linear Models: Precise Asymptotics in High-dimensionsBruno Loureiro, Gabriele Sicuro, Cédric Gerbelot, Alessandro Pacco 等NeurIPS 2021 · 被引用 70 次
