Learnability of convolutional neural networks for infinite dimensional input via mixed and anisotropic smoothness
Sho Okumoto, Taiji Suzuki
Abstract
Among a wide range of success of deep learning, convolutional neural networks have been extensively utilized in several tasks such as speech recognition, image processing, and natural language processing, which require inputs with large dimensions.Several studies have investigated function estimation capability of deep learning, but most of them have assumed that the dimensionality of the input is much smaller than the sample size. However, for typical data in applications such as those handled by the convolutional neural networks described above, the dimensionality of inputs is relatively high or even infinite. In this paper, we investigate the approximation and estimation errors of the (dilated) convolutional neural networks when the input is infinite dimensional. Although the approximation and estimation errors of neural networks are affected by the curse of dimensionality in the existing analyses for typical function spaces such as the and Besov spaces, we show that, by considering anisotropic smoothness, they can alleviate exponential dependency on the dimensionality but they only depend on the smoothness of the target functions. Our theoretical analysis supports the great practical success of convolutional networks. Furthermore, we show that the dilated convolution is advantageous when the smoothness of the target function has a sparse structure.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers8
- Transformers are Minimax Optimal Nonparametric In-Context LearnersJuno Kim, Tai Nakamaki, Taiji SuzukiNeurIPS 2024 · 42 citations
- Approximation and Estimation Ability of Transformers for Sequence-to-Sequence Functions with Infinite Dimensional InputShokichi Takakura, Taiji SuzukiICML 2023 · 32 citations
- Universality of Group Convolutional Neural Networks Based on Ridgelet Analysis on GroupsSho Sonoda, Isao Ishikawa, Masahiro IkedaNeurIPS 2022 · 12 citations
- Theoretical Analysis of the Inductive Biases in Deep Convolutional NetworksZihao Wang, Lei WuNeurIPS 2023 · 10 citations
- Minimax optimality of convolutional neural networks for infinite dimensional input-output problems and separation from kernel methodsYuto Nishimura, Taiji SuzukiICLR 2024 · 3 citations
Related papers
- Deep learning is adaptive to intrinsic dimensionality of model smoothness in anisotropic Besov spaceTaiji Suzuki, Atsushi NitandaNeurIPS 2021 · 76 citations
- Posterior Contraction for Sparse Neural Networks in Besov Spaces with Intrinsic DimensionalityKyeongwon Lee, Lizhen Lin, Jaewoo Park, Seonghyun JeongNeurIPS 2025 · 4 citations
- What Can Be Learnt With Wide Convolutional Neural Networks?Francesco Cagnetta, Alessandro Favero, Matthieu WyartICML 2023 · 16 citations
- Approximation and non-parametric estimation of functions over high-dimensional spheres via deep ReLU networksNamjoon Suh, Tian-Yi Zhou, Xiaoming HuoICLR 2023
- Deep Learning meets Nonparametric Regression: Are Weight-Decayed DNNs Locally Adaptive?Kaiqi Zhang, Yu-Xiang WangICLR 2023 · 3 citations
