Characterization of Gaussian Universality Breakdown in High-Dimensional Empirical Risk Minimization
Mohamed Chiheb Yaakoubi, Cosme Louart, Malik TIOMOKO, Zhenyu Liao
摘要
We study high-dimensional convex empirical risk minimization (ERM) under general non-Gaussian data designs. By heuristically extending the Convex Gaussian Min–Max Theorem (CGMT) to non-Gaussian settings, we derive an asymptotic min–max characterization of key statistics, enabling approximation of the mean and covariance of the ERM estimator . Specifically, under a concentration assumption on the data matrix and standard regularity conditions on the loss and regularizer, we show that for a test covariate independent of the training data, the projection approximately follows the convolution of the (generally non-Gaussian) distribution of with an independent centered Gaussian variable of variance . This result clarifies the scope and limits of Gaussian universality for ERMs. Additionally, we prove that any regularizer is asymptotically equivalent to a quadratic form determined solely by its Hessian at zero and gradient at . Numerical simulations across diverse losses and models are provided to validate our theoretical predictions and qualitative insights.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Learning curves of generic features maps for realistic datasets with a teacher-student modelBruno Loureiro, Cédric Gerbelot, Hugo Cui, Sebastian Goldt 等NeurIPS 2021 · 被引用 170 次
- Random Matrix Theory Proves that Deep Learning Representations of GAN-data Behave as Gaussian MixturesMohamed El Amine Seddik, Cosme Louart, Mohamed Tamaazousti, Romain CouilletICML 2020 · 被引用 78 次
- Learning Gaussian Mixtures with Generalized Linear Models: Precise Asymptotics in High-dimensionsBruno Loureiro, Gabriele Sicuro, Cédric Gerbelot, Alessandro Pacco 等NeurIPS 2021 · 被引用 70 次
- Universality laws for Gaussian mixtures in generalized linear modelsYatin Dandi, Ludovic Stephan, Florent Krzakala, Bruno Loureiro 等NeurIPS 2023 · 被引用 40 次
- Are Gaussian Data All You Need? The Extents and Limits of Universality in High-Dimensional Generalized Linear EstimationLuca Pesce, Florent Krzakala, Bruno Loureiro, Ludovic StephanICML 2023 · 被引用 6 次
相关 Paper
- The Breakdown of Gaussian Universality in Classification of High-dimensional Linear Factor MixturesXiaoyi Mai, Zhenyu LiaoICLR 2025
- Classification of Heavy-tailed Features in High Dimensions: a Superstatistical ApproachUrte Adomaityte, Gabriele Sicuro, Pierpaolo VivoNeurIPS 2023 · 被引用 17 次
- On the Asymptotic Distribution of the Minimum Empirical RiskJacob Westerhout, TrungTin Nguyen, Xin Guo, Hien Duy NguyenICML 2024 · 被引用 8 次
- The Nuclear Route: Sharp Asymptotics of ERM in Overparameterized Quadratic NetworksVittorio Erba, Emanuele Troiani, Lenka Zdeborová, Florent KrzakalaNeurIPS 2025 · 被引用 13 次
- Optimal Excess Risk Bounds for Empirical Risk Minimization on p-Norm Linear RegressionAyoub El Hanchi, Murat A. ErdogduNeurIPS 2023 · 被引用 2 次
