Minimax optimality of convolutional neural networks for infinite dimensional input-output problems and separation from kernel methods
Yuto Nishimura, Taiji Suzuki
Abstract
Recent deep learning applications, exemplified by text-to-image tasks, often involve high-dimensional inputs and outputs. While several studies have investigated the function estimation capabilities of deep learning, research on dilated convolutional neural networks (CNNs) has mainly focused on cases where input dimensions are infinite but output dimensions are one-dimensional, similar to many other studies. However, many practical deep learning tasks involve highdimensional (or even infinite dimensional) inputs and outputs. In this paper, we investigate the optimality of dilated CNNs for estimating a map between infinitedimensional input and output spaces by analyzing their approximation and estimation abilities. For that purpose, we first show that approximation and estimation errors depend only on the smoothness and decay rate with respect to the infinity norm of the output, and their estimation accuracy actually achieve the minimax optimal rate of convergence. Second, we demonstrate that the dilated CNNs outperform any linear estimators including kernel ridge regression and k-NN estimators in a minimax error sense, highlighting the usefulness of feature learning realized by deep neural networks. Our theoretical analysis provide a theoretical basis for understanding the success of deep learning in recent high-dimensional input-output tasks. 1. We consider a nonlinear operator as a true function with infinite-dimensional inputs and outputs. In the aforementioned setup, we demonstrate that dilated CNNs achieve approximation and estimation errors that depend only on the smoothness and decay rate of the output. Furthermore,
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Transformers are Minimax Optimal Nonparametric In-Context LearnersJuno Kim, Tai Nakamaki, Taiji SuzukiNeurIPS 2024 · 42 citations
- Meta Optimality for Demographic Parity Constrained Regression via Post-ProcessingKazuto FukuchiICML 2025
Builds on8
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- Fourier Neural Operator for Parametric Partial Differential EquationsZongyi Li, Nikola Borislavov Kovachki, Kamyar Azizzadenesheli, Burigede Liu et al.ICLR 2021 · 3,911 citations
- Minimax Estimation of Conditional Moment ModelsNishanth Dikkala, Greg Lewis, Lester Mackey, Vasilis SyrgkanisNeurIPS 2020 · 125 citations
- Dual Instrumental Variable RegressionKrikamol Muandet, Arash Mehrjou, Si Kai Lee, Anant RajNeurIPS 2020 · 87 citations
Related papers
- Learnability of convolutional neural networks for infinite dimensional input via mixed and anisotropic smoothnessSho Okumoto, Taiji SuzukiICLR 2022 · 9 citations
- Deep Neural Network Regression with Functional CovariatesHang Zhou, Ju-Sheng Hong, Xiucai Ding, Jane-Ling WangICML 2026
- Deep learning is adaptive to intrinsic dimensionality of model smoothness in anisotropic Besov spaceTaiji Suzuki, Atsushi NitandaNeurIPS 2021 · 76 citations
- Approximation and Estimation Ability of Transformers for Sequence-to-Sequence Functions with Infinite Dimensional InputShokichi Takakura, Taiji SuzukiICML 2023 · 32 citations
- What Can Be Learnt With Wide Convolutional Neural Networks?Francesco Cagnetta, Alessandro Favero, Matthieu WyartICML 2023 · 16 citations
