Learning from higher-order correlations, efficiently: hypothesis tests, random features, and neural networks
Eszter Székely, Lorenzo Bardone, Federica Gerace, Sebastian Goldt
摘要
Neural networks excel at discovering statistical patterns in high-dimensional data sets. In practice, higher-order cumulants, which quantify the non-Gaussian correlations between three or more variables, are particularly important for the performance of neural networks. But how efficient are neural networks at extracting features from higher-order cumulants? We study this question in the spiked cumulant model, where the statistician needs to recover a privileged direction or ‘spike’ from the order- p⩾4 cumulants of d-dimensional inputs. We first discuss the fundamental statistical and computational limits of recovering the spike by analysing the number of samples n required to strongly distinguish between inputs from the spiked cumulant model and isotropic Gaussian inputs. Existing literature established the presence of a wide statistical-to-computational gap in this problem. We deepen this line of work by finding an exact formula for the likelihood ratio norm which proves that statistical distinguishability requires n≳d samples, while distinguishing the two distributions in polynomial time requires n≳d2 samples for a wide class of algorithms, i.e. those covered by the low-degree conjecture. Numerical experiments show that neural networks do indeed learn to distinguish the two distributions with quadratic sample complexity, while ‘lazy’ methods like random features (RFs) are not better than random guessing in this regime. Our results show that neural networks extract information from higher-order correlations in the spiked cumulant model efficiently, and reveal a large gap in the amount of data required by neural networks and RFs to learn from higher-order cumulants.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- A Solvable High-Dimensional Model Where Nonlinear Autoencoders Learn Structure Invisible to PCA While Test Loss Misaligns With GeneralizationVicente Mendes, Lorenzo Bardone, Cédric Koller, Jorge Medina Moreira 等ICML 2026 · 被引用 6 次
- A theory of learning data statistics in diffusion models, from easy to hardLorenzo Bardone, Claudia Merger, Sebastian GoldtICML 2026
- LeSTD: LLM Compression via Learning-based Sparse Tensor DecompositionYi Li, Zhichun Guo, Miao Yin, Bingzhe LiICLR 2026
它引用的顶会 Paper4
- Learning Parities with Neural NetworksAmit Daniely, Eran MalachNeurIPS 2020 · 被引用 104 次
- Learning Gaussian Mixtures with Generalized Linear Models: Precise Asymptotics in High-dimensionsBruno Loureiro, Gabriele Sicuro, Cédric Gerbelot, Alessandro Pacco 等NeurIPS 2021 · 被引用 70 次
- Precise Learning Curves and Higher-Order Scalings for Dot-product Kernel RegressionLechao Xiao, Hong Hu, Theodor Misiakiewicz, Yue Lu 等NeurIPS 2022 · 被引用 27 次
- Sliding Down the Stairs: How Correlated Latent Variables Accelerate Learning with Neural NetworksLorenzo Bardone, Sebastian GoldtICML 2024 · 被引用 13 次
相关 Paper
- Learning in the Presence of Low-dimensional Structure: A Spiked Random Matrix PerspectiveJimmy Ba, Murat A. Erdogdu, Taiji Suzuki, Zhichao Wang 等NeurIPS 2023 · 被引用 47 次
- Computational barriers for permutation-based problems, and cumulants of weakly dependent random variablesBertrand Even, Christophe Giraud, Nicolas VerzelenSODA 2026
- Identification of Causal Structure with Latent Variables Based on Higher Order CumulantsWei Chen, Zhiyi Huang, Ruichu Cai, Zhifeng Hao 等AAAI 2024 · 被引用 10 次
- When Do Neural Networks Outperform Kernel Methods?Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, Andrea MontanariNeurIPS 2020 · 被引用 217 次
- Tensor Cumulants for Statistical Inference on Invariant DistributionsDmitriy Kunisky, Cristopher Moore, Alexander S. WeinFOCS 2024 · 被引用 7 次
