What Can Be Learnt With Wide Convolutional Neural Networks?
Francesco Cagnetta, Alessandro Favero, Matthieu Wyart
摘要
Understanding how convolutional neural networks (CNNs) can efficiently learn high-dimensional functions remains a fundamental challenge. A popular belief is that these models harness the local and hierarchical structure of natural data such as images. Yet, we lack a quantitative understanding of how such structure affects performance, for example the rate of decay of the generalisation error with the number of training samples. In this paper, we study infinitely wide deep CNNs in the kernel regime. First, we show that the spectrum of the corresponding kernel inherits the hierarchical structure of the network, and we characterise its asymptotics. Then, we use this result together with generalisation bounds to prove that deep CNNs adapt to the spatial scale of the target function. In particular, we find that if the target function depends on low-dimensional subsets of adjacent input variables then the decay of the error is controlled by the effective dimensionality of these subsets. Conversely, if the target function depends on the full set of input variables then the error decay is controlled by the input dimension. We conclude by computing the generalisation error of a deep CNN trained on the output of another deep CNN with randomly initialised parameters. Interestingly, we find that, despite their hierarchical structure, the functions generated by infinitely wide deep CNNs are too rich to be efficiently learnable in high dimensions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Task Arithmetic in the Tangent Space: Improved Editing of Pre-Trained ModelsGuillermo Ortiz-Jiménez, Alessandro Favero, Pascal FrossardNeurIPS 2023 · 被引用 272 次
- Towards a theory of how the structure of language is acquired by deep neural networksFrancesco Cagnetta, Matthieu WyartNeurIPS 2024 · 被引用 33 次
- Dimension-free deterministic equivalents and scaling laws for random feature regressionLeonardo Defilippis, Bruno Loureiro, Theodor MisiakiewiczNeurIPS 2024 · 被引用 28 次
- Towards Understanding Inductive Bias in Transformers: A View From InfinityItay Lavie, Guy Gur-Ari, Zohar RingelICML 2024 · 被引用 11 次
- How Deep Networks Learn Sparse and Hierarchical Data: the Sparse Random Hierarchy ModelUmberto M. Tomasini, Matthieu WyartICML 2024 · 被引用 7 次
它引用的顶会 Paper15
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan 等ICLR 2020 · 被引用 705 次
- Spectrum Dependent Learning Curves in Kernel Regression and Wide Neural NetworksBlake Bordelon, Abdulkadir Canatar, Cengiz PehlevanICML 2020 · 被引用 245 次
- Learning curves of generic features maps for realistic datasets with a teacher-student modelBruno Loureiro, Cédric Gerbelot, Hugo Cui, Sebastian Goldt 等NeurIPS 2021 · 被引用 170 次
- Generalization Error Rates in Kernel Regression: The Crossover from the Noiseless to Noisy RegimeHugo Cui, Bruno Loureiro, Florent Krzakala, Lenka ZdeborováNeurIPS 2021 · 被引用 109 次
- Implicit Regularization of Random Feature ModelsArthur Jacot, Berfin Simsek, Francesco Spadaro, Clément Hongler 等ICML 2020 · 被引用 83 次
相关 Paper
- Approximation and Learning with Deep Convolutional Models: a Kernel PerspectiveAlberto BiettiICLR 2022 · 被引用 33 次
- Learnability of convolutional neural networks for infinite dimensional input via mixed and anisotropic smoothnessSho Okumoto, Taiji SuzukiICLR 2022 · 被引用 9 次
- Disentangling Trainability and Generalization in Deep Neural NetworksLechao Xiao, Jeffrey Pennington, Samuel Stern SchoenholzICML 2020 · 被引用 91 次
- Learning with convolution and pooling operations in kernel methodsTheodor Misiakiewicz, Song MeiNeurIPS 2022 · 被引用 30 次
- On the Spectral Bias of Convolutional Neural Tangent and Gaussian Process KernelsAmnon Geifman, Meirav Galun, David Jacobs, Ronen BasriNeurIPS 2022 · 被引用 22 次
