Locality defeats the curse of dimensionality in convolutional teacher-student scenarios
Alessandro Favero, Francesco Cagnetta, Matthieu Wyart
摘要
Convolutional neural networks perform a local and translationally-invariant treatment of the data: quantifying which of these two aspects is central to their success remains a challenge. We study this problem within a teacher–student framework for kernel regression, using ‘convolutional’ kernels inspired by the neural tangent kernel of simple convolutional architectures of given filter size. Using heuristic methods from physics, we find in the ridgeless case that locality is key in determining the learning curve exponent β (that relates the test error ϵ t ∼ P −β to the size of the training set P), whereas translational invariance is not. In particular, if the filter size of the teacher t is smaller than that of the student s, β is a function of s only and does not depend on the input dimension. We confirm our predictions on β empirically. We conclude by proving, under a natural universality assumption, that performing kernel regression with a ridge that decreases with the size of the training set leads to similar learning curve exponents to those we obtain in the ridgeless case.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Task Arithmetic in the Tangent Space: Improved Editing of Pre-Trained ModelsGuillermo Ortiz-Jiménez, Alessandro Favero, Pascal FrossardNeurIPS 2023 · 被引用 272 次
- Short-Term Memory ConvolutionsGrzegorz Stefanski, Krzysztof Arendt, Pawel Daniluk, Bartlomiej Jasik 等ICLR 2023 · 被引用 67 次
- The SSL Interplay: Augmentations, Inductive Bias, and GeneralizationVivien Cabannes, Bobak Toussi Kiani, Randall Balestriero, Yann LeCun 等ICML 2023 · 被引用 43 次
- Approximation and Learning with Deep Convolutional Models: a Kernel PerspectiveAlberto BiettiICLR 2022 · 被引用 33 次
- Dimension-free deterministic equivalents and scaling laws for random feature regressionLeonardo Defilippis, Bruno Loureiro, Theodor MisiakiewiczNeurIPS 2024 · 被引用 28 次
它引用的顶会 Paper6
- Spectrum Dependent Learning Curves in Kernel Regression and Wide Neural NetworksBlake Bordelon, Abdulkadir Canatar, Cengiz PehlevanICML 2020 · 被引用 245 次
- On the Similarity between the Laplace and Neural Tangent KernelsAmnon Geifman, Abhay Kumar Yadav, Yoni Kasten, Meirav Galun 等NeurIPS 2020 · 被引用 118 次
- Implicit Regularization of Random Feature ModelsArthur Jacot, Berfin Simsek, Francesco Spadaro, Clément Hongler 等ICML 2020 · 被引用 83 次
- Towards Learning Convolutions from ScratchBehnam NeyshaburNeurIPS 2020 · 被引用 80 次
- Kernel Alignment Risk Estimator: Risk Prediction from Training DataArthur Jacot, Berfin Simsek, Francesco Spadaro, Clément Hongler 等NeurIPS 2020 · 被引用 74 次
相关 Paper
- On the Asymptotic Learning Curves of Kernel Ridge Regression under Power-law DecayYicheng Li, Haobo Zhang, Qian LinNeurIPS 2023 · 被引用 23 次
- Excess Risk of Two-Layer ReLU Neural Networks in Teacher-Student Settings and its Superiority to Kernel MethodsShunta Akiyama, Taiji SuzukiICLR 2023 · 被引用 1 次
- High-dimensional Asymptotics of Feature Learning: How One Gradient Step Improves the RepresentationJimmy Ba, Murat A. Erdogdu, Taiji Suzuki, Zhichao Wang 等NeurIPS 2022 · 被引用 173 次
- Learning Curves for Gaussian Process Regression with Power-Law Priors and TargetsHui Jin, Pradeep Kr. Banerjee, Guido MontúfarICLR 2022 · 被引用 18 次
- The Surprising Simplicity of the Early-Time Learning Dynamics of Neural NetworksWei Hu, Lechao Xiao, Ben Adlam, Jeffrey PenningtonNeurIPS 2020 · 被引用 77 次
