Implicit Regularization of Random Feature Models
Arthur Jacot, Berfin Simsek, Francesco Spadaro, Clément Hongler, Franck Gabriel
Abstract
Random Feature (RF) models are used as efficient parametric approximations of kernel methods. We investigate, by means of random matrix theory, the connection between Gaussian RF models and Kernel Ridge Regression (KRR). For a Gaussian RF model with features, data points, and a ridge , we show that the average (i.e. expected) RF predictor is close to a KRR predictor with an effective ridge . We show that and monotonically as grows, thus revealing the implicit regularization effect of finite RF sampling. We then compare the risk (i.e. test error) of the -KRR predictor with the average risk of the -RF predictor and obtain a precise and explicit bound on their difference. Finally, we empirically find an extremely good agreement between the test errors of the average -RF predictor and -KRR predictor.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 357cf93e-511c-472f-949c-3454df662464Cited by top-tier papers33
- Understanding Double Descent Requires A Fine-Grained Bias-Variance DecompositionBen Adlam, Jeffrey PenningtonNeurIPS 2020 · 111 citations
- More Than a Toy: Random Matrix Models Predict How Real-World Neural Representations GeneralizeAlexander Wei, Wei Hu, Jacob SteinhardtICML 2022 · 90 citations
- Kernel Alignment Risk Estimator: Risk Prediction from Training DataArthur Jacot, Berfin Simsek, Francesco Spadaro, Clément Hongler et al.NeurIPS 2020 · 74 citations
- Early Stopping in Deep Networks: Double Descent and How to Eliminate itReinhard Heckel, Fatih Furkan YilmazICLR 2021 · 55 citations
- Scaling Neural Tangent Kernels via Sketching and Random FeaturesAmir Zandieh, Insu Han, Haim Avron, Neta Shoham et al.NeurIPS 2021 · 42 citations
Builds on3
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang et al.ICLR 2020 · 1,108 citations
- Double Trouble in Double Descent: Bias and Variance(s) in the Lazy RegimeStéphane d'Ascoli, Maria Refinetti, Giulio Biroli, Florent KrzakalaICML 2020 · 163 citations
- Ridge Regression: Structure, Cross-Validation, and SketchingSifan Liu, Edgar DobribanICLR 2020 · 52 citations
Related papers
- Dimension-free deterministic equivalents and scaling laws for random feature regressionLeonardo Defilippis, Bruno Loureiro, Theodor MisiakiewiczNeurIPS 2024 · 28 citations
- Error Bounds for Learning with Vector-Valued Random FeaturesSamuel Lanthaler, Nicholas H. NelsenNeurIPS 2023 · 26 citations
- More is Better: when Infinite Overparameterization is Optimal and Overfitting is ObligatoryJames B. Simon, Dhruva Karkada, Nikhil Ghosh, Mikhail BelkinICLR 2024 · 7 citations
- Quasi-Monte Carlo Features for Kernel ApproximationZhen Huang, Jiajin Sun, Yian HuangICML 2024 · 6 citations
- On the Inherent Regularization Effects of Noise Injection During TrainingOussama Dhifallah, Yue M. LuICML 2021 · 36 citations
