Characterization of Gaussian Universality Breakdown in High-Dimensional Empirical Risk Minimization
Mohamed Chiheb Yaakoubi, Cosme Louart, Malik TIOMOKO, Zhenyu Liao
Abstract
We study high-dimensional convex empirical risk minimization (ERM) under general non-Gaussian data designs. By heuristically extending the Convex Gaussian Min–Max Theorem (CGMT) to non-Gaussian settings, we derive an asymptotic min–max characterization of key statistics, enabling approximation of the mean and covariance of the ERM estimator . Specifically, under a concentration assumption on the data matrix and standard regularity conditions on the loss and regularizer, we show that for a test covariate independent of the training data, the projection approximately follows the convolution of the (generally non-Gaussian) distribution of with an independent centered Gaussian variable of variance . This result clarifies the scope and limits of Gaussian universality for ERMs. Additionally, we prove that any regularizer is asymptotically equivalent to a quadratic form determined solely by its Hessian at zero and gradient at . Numerical simulations across diverse losses and models are provided to validate our theoretical predictions and qualitative insights.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 545a7594-4ddd-46cb-b915-bd8d3816f9f8Builds on6
- Learning curves of generic features maps for realistic datasets with a teacher-student modelBruno Loureiro, Cédric Gerbelot, Hugo Cui, Sebastian Goldt et al.NeurIPS 2021 · 170 citations
- Random Matrix Theory Proves that Deep Learning Representations of GAN-data Behave as Gaussian MixturesMohamed El Amine Seddik, Cosme Louart, Mohamed Tamaazousti, Romain CouilletICML 2020 · 78 citations
- Learning Gaussian Mixtures with Generalized Linear Models: Precise Asymptotics in High-dimensionsBruno Loureiro, Gabriele Sicuro, Cédric Gerbelot, Alessandro Pacco et al.NeurIPS 2021 · 70 citations
- Universality laws for Gaussian mixtures in generalized linear modelsYatin Dandi, Ludovic Stephan, Florent Krzakala, Bruno Loureiro et al.NeurIPS 2023 · 40 citations
- Are Gaussian Data All You Need? The Extents and Limits of Universality in High-Dimensional Generalized Linear EstimationLuca Pesce, Florent Krzakala, Bruno Loureiro, Ludovic StephanICML 2023 · 6 citations
Related papers
- The Breakdown of Gaussian Universality in Classification of High-dimensional Linear Factor MixturesXiaoyi Mai, Zhenyu LiaoICLR 2025
- Classification of Heavy-tailed Features in High Dimensions: a Superstatistical ApproachUrte Adomaityte, Gabriele Sicuro, Pierpaolo VivoNeurIPS 2023 · 17 citations
- On the Asymptotic Distribution of the Minimum Empirical RiskJacob Westerhout, TrungTin Nguyen, Xin Guo, Hien Duy NguyenICML 2024 · 8 citations
- The Nuclear Route: Sharp Asymptotics of ERM in Overparameterized Quadratic NetworksVittorio Erba, Emanuele Troiani, Lenka Zdeborová, Florent KrzakalaNeurIPS 2025 · 13 citations
- Optimal Excess Risk Bounds for Empirical Risk Minimization on p-Norm Linear RegressionAyoub El Hanchi, Murat A. ErdogduNeurIPS 2023 · 2 citations
