Predicting Kernel Regression Learning Curves from Only Raw Data Statistics
Dhruva Karkada, Joseph Turnbull, Yuxi Liu, James B Simon
Abstract
We study kernel regression with common rotation-invariant kernels on real datasets including CIFAR-5m, SVHN, and ImageNet. We give a theoretical framework that predicts learning curves (test risk vs. sample size) from only two measurements: the empirical data covariance matrix and an empirical polynomial decomposition of the target function f * . The key new idea is an analytical approximation of a kernel's eigenvalues and eigenfunctions with respect to an anisotropic data distribution. The eigenfunctions resemble Hermite polynomials of the data, so we call this approximation the Hermite eigenstructure ansatz (HEA). We prove the HEA for Gaussian data, but we find that real image data is often "Gaussian enough" for the HEA to hold well in practice, enabling us to predict learning curves by applying prior results relating kernel eigenstructure to test risk. Extending beyond kernel regression, we empirically find that MLPs in the feature-learning regime learn Hermite polynomials in the order predicted by the HEA. Our HEA framework is a proof of concept that an end-to-end theory of learning which maps dataset structure all the way to model performance is possible for nontrivial learning algorithms on real datasets. • UC Berkeley Imbue ⋆ Joint primary authorship. Work completed during summer internship at Imbue. Code: https://github.com/JoeyTurn/hermite-eigenstructure-ansatz 1 See Section 3.2 and Section A for a review of Hermite polynomials.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dc5e348d-54ea-48f6-b68f-89dca649886aBuilds on14
- Spectrum Dependent Learning Curves in Kernel Regression and Wide Neural NetworksBlake Bordelon, Abdulkadir Canatar, Cengiz PehlevanICML 2020 · 245 citations
- Tensor Programs IV: Feature Learning in Infinite-Width Neural NetworksGreg Yang, Edward J. HuICML 2021 · 242 citations
- When Do Neural Networks Outperform Kernel Methods?Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, Andrea MontanariNeurIPS 2020 · 217 citations
- Learning curves of generic features maps for realistic datasets with a teacher-student modelBruno Loureiro, Cédric Gerbelot, Hugo Cui, Sebastian Goldt et al.NeurIPS 2021 · 170 citations
- Learning single-index models with shallow neural networksAlberto Bietti, Joan Bruna, Clayton Sanford, Min Jae SongNeurIPS 2022 · 119 citations
Related papers
- Kernel Alignment Risk Estimator: Risk Prediction from Training DataArthur Jacot, Berfin Simsek, Francesco Spadaro, Clément Hongler et al.NeurIPS 2020 · 74 citations
- Precise Learning Curves and Higher-Order Scalings for Dot-product Kernel RegressionLechao Xiao, Hong Hu, Theodor Misiakiewicz, Yue Lu et al.NeurIPS 2022 · 27 citations
- A Comprehensive Analysis on the Learning Curve in Kernel Ridge RegressionTin Sum Cheng, Aurélien Lucchi, Anastasis Kratsios, David BeliusNeurIPS 2024 · 7 citations
- Learning Curves for Gaussian Process Regression with Power-Law Priors and TargetsHui Jin, Pradeep Kr. Banerjee, Guido MontúfarICLR 2022 · 18 citations
- Locality defeats the curse of dimensionality in convolutional teacher-student scenariosAlessandro Favero, Francesco Cagnetta, Matthieu WyartNeurIPS 2021 · 34 citations
