Minimax optimality of convolutional neural networks for infinite dimensional input-output problems and separation from kernel methods
Yuto Nishimura, Taiji Suzuki
摘要
Recent deep learning applications, exemplified by text-to-image tasks, often involve high-dimensional inputs and outputs. While several studies have investigated the function estimation capabilities of deep learning, research on dilated convolutional neural networks (CNNs) has mainly focused on cases where input dimensions are infinite but output dimensions are one-dimensional, similar to many other studies. However, many practical deep learning tasks involve highdimensional (or even infinite dimensional) inputs and outputs. In this paper, we investigate the optimality of dilated CNNs for estimating a map between infinitedimensional input and output spaces by analyzing their approximation and estimation abilities. For that purpose, we first show that approximation and estimation errors depend only on the smoothness and decay rate with respect to the infinity norm of the output, and their estimation accuracy actually achieve the minimax optimal rate of convergence. Second, we demonstrate that the dilated CNNs outperform any linear estimators including kernel ridge regression and k-NN estimators in a minimax error sense, highlighting the usefulness of feature learning realized by deep neural networks. Our theoretical analysis provide a theoretical basis for understanding the success of deep learning in recent high-dimensional input-output tasks. 1. We consider a nonlinear operator as a true function with infinite-dimensional inputs and outputs. In the aforementioned setup, we demonstrate that dilated CNNs achieve approximation and estimation errors that depend only on the smoothness and decay rate of the output. Furthermore,
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Transformers are Minimax Optimal Nonparametric In-Context LearnersJuno Kim, Tai Nakamaki, Taiji SuzukiNeurIPS 2024 · 被引用 42 次
- Meta Optimality for Demographic Parity Constrained Regression via Post-ProcessingKazuto FukuchiICML 2025
它引用的顶会 Paper8
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- Fourier Neural Operator for Parametric Partial Differential EquationsZongyi Li, Nikola Borislavov Kovachki, Kamyar Azizzadenesheli, Burigede Liu 等ICLR 2021 · 被引用 3,911 次
- Minimax Estimation of Conditional Moment ModelsNishanth Dikkala, Greg Lewis, Lester Mackey, Vasilis SyrgkanisNeurIPS 2020 · 被引用 125 次
- Dual Instrumental Variable RegressionKrikamol Muandet, Arash Mehrjou, Si Kai Lee, Anant RajNeurIPS 2020 · 被引用 87 次
相关 Paper
- Learnability of convolutional neural networks for infinite dimensional input via mixed and anisotropic smoothnessSho Okumoto, Taiji SuzukiICLR 2022 · 被引用 9 次
- Deep Neural Network Regression with Functional CovariatesHang Zhou, Ju-Sheng Hong, Xiucai Ding, Jane-Ling WangICML 2026
- Deep learning is adaptive to intrinsic dimensionality of model smoothness in anisotropic Besov spaceTaiji Suzuki, Atsushi NitandaNeurIPS 2021 · 被引用 76 次
- Approximation and Estimation Ability of Transformers for Sequence-to-Sequence Functions with Infinite Dimensional InputShokichi Takakura, Taiji SuzukiICML 2023 · 被引用 32 次
- What Can Be Learnt With Wide Convolutional Neural Networks?Francesco Cagnetta, Alessandro Favero, Matthieu WyartICML 2023 · 被引用 16 次
