Implicit Regularization of Random Feature Models
Arthur Jacot, Berfin Simsek, Francesco Spadaro, Clément Hongler, Franck Gabriel
摘要
Random Feature (RF) models are used as efficient parametric approximations of kernel methods. We investigate, by means of random matrix theory, the connection between Gaussian RF models and Kernel Ridge Regression (KRR). For a Gaussian RF model with features, data points, and a ridge , we show that the average (i.e. expected) RF predictor is close to a KRR predictor with an effective ridge . We show that and monotonically as grows, thus revealing the implicit regularization effect of finite RF sampling. We then compare the risk (i.e. test error) of the -KRR predictor with the average risk of the -RF predictor and obtain a precise and explicit bound on their difference. Finally, we empirically find an extremely good agreement between the test errors of the average -RF predictor and -KRR predictor.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper33
- Understanding Double Descent Requires A Fine-Grained Bias-Variance DecompositionBen Adlam, Jeffrey PenningtonNeurIPS 2020 · 被引用 111 次
- More Than a Toy: Random Matrix Models Predict How Real-World Neural Representations GeneralizeAlexander Wei, Wei Hu, Jacob SteinhardtICML 2022 · 被引用 90 次
- Kernel Alignment Risk Estimator: Risk Prediction from Training DataArthur Jacot, Berfin Simsek, Francesco Spadaro, Clément Hongler 等NeurIPS 2020 · 被引用 74 次
- Early Stopping in Deep Networks: Double Descent and How to Eliminate itReinhard Heckel, Fatih Furkan YilmazICLR 2021 · 被引用 55 次
- Scaling Neural Tangent Kernels via Sketching and Random FeaturesAmir Zandieh, Insu Han, Haim Avron, Neta Shoham 等NeurIPS 2021 · 被引用 42 次
它引用的顶会 Paper3
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang 等ICLR 2020 · 被引用 1,108 次
- Double Trouble in Double Descent: Bias and Variance(s) in the Lazy RegimeStéphane d'Ascoli, Maria Refinetti, Giulio Biroli, Florent KrzakalaICML 2020 · 被引用 163 次
- Ridge Regression: Structure, Cross-Validation, and SketchingSifan Liu, Edgar DobribanICLR 2020 · 被引用 52 次
相关 Paper
- Dimension-free deterministic equivalents and scaling laws for random feature regressionLeonardo Defilippis, Bruno Loureiro, Theodor MisiakiewiczNeurIPS 2024 · 被引用 28 次
- Error Bounds for Learning with Vector-Valued Random FeaturesSamuel Lanthaler, Nicholas H. NelsenNeurIPS 2023 · 被引用 26 次
- More is Better: when Infinite Overparameterization is Optimal and Overfitting is ObligatoryJames B. Simon, Dhruva Karkada, Nikhil Ghosh, Mikhail BelkinICLR 2024 · 被引用 7 次
- Quasi-Monte Carlo Features for Kernel ApproximationZhen Huang, Jiajin Sun, Yian HuangICML 2024 · 被引用 6 次
- On the Inherent Regularization Effects of Noise Injection During TrainingOussama Dhifallah, Yue M. LuICML 2021 · 被引用 36 次
