Theoretical Analysis of the Inductive Biases in Deep Convolutional Networks
Zihao Wang, Lei Wu
摘要
In this paper, we provide a theoretical analysis of the inductive biases in convolutional neural networks (CNNs). We start by examining the universality of CNNs, i.e., the ability to approximate any continuous functions. We prove that a depth of O(log d) suffices for deep CNNs to achieve this universality, where d in the input dimension. Additionally, we establish that learning sparse functions with CNNs requires only O(log 2 d) samples, indicating that deep CNNs can efficiently capture long-range sparse correlations. These results are made possible through a novel combination of the multichanneling and downsampling when increasing the network depth. We also delve into the distinct roles of weight sharing and locality in CNNs. To this end, we compare the performance of CNNs, locally-connected networks (LCNs), and fully-connected networks (FCNs) on a simple regression task, where LCNs can be viewed as CNNs without weight sharing. On the one hand, we prove that LCNs require Ω(d) samples while CNNs need only O(log 2 d) samples, highlighting the critical role of weight sharing. On the other hand, we prove that FCNs require Ω(d 2 ) samples, whereas LCNs need only O(d) samples, underscoring the importance of locality. These provable separations quantify the difference between the two biases, and the major observation behind our proof is that weight sharing and locality break different symmetries in the learning process.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- On the Inductive Bias of Stacking Towards Improving ReasoningNikunj Saunshi, Stefani Karp, Shankar Krishnan, Sobhan Miryoosefi 等NeurIPS 2024 · 被引用 23 次
- Which Frequencies do CNNs Need? Emergent Bottleneck Structure in Feature LearningYuxiao Wen, Arthur JacotICML 2024 · 被引用 9 次
- Separation and Bias of Deep Equilibrium Models on Expressivity and Learning DynamicsZhoutong Wu, Yimu Zhang, Cong Fang, Zhouchen LinNeurIPS 2024 · 被引用 3 次
- EasySpiro: Assessing Lung Function via Arbitrary Exhalations on Commodity EarphonesChi Xu, Wentao Xie, Baichen Yang, Yizhen Zhang 等MobiCom 2025 · 被引用 3 次
- From 2D Grids to 1D Tokens: Reforming Shared Representations for Multimodal Image FusionYuchen Xian, Yunqiu Xu, Yang He, Yi YangICML 2026 · 被引用 2 次
它引用的顶会 Paper11
- Searching for Efficient Transformers for Language ModelingDavid R. So, Wojciech Manke, Hanxiao Liu, Zihang Dai 等NeurIPS 2021 · 被引用 205 次
- Generalization bounds for deep convolutional neural networksPhilip M. Long, Hanie SedghiICLR 2020 · 被引用 102 次
- On the Sample Complexity of Learning under Geometric StabilityAlberto Bietti, Luca Venturi, Joan BrunaNeurIPS 2021 · 被引用 45 次
- Implicit Regularization in Hierarchical Tensor Factorization and Deep Convolutional Neural NetworksNoam Razin, Asaf Maman, Nadav CohenICML 2022 · 被引用 34 次
- Approximation and Learning with Deep Convolutional Models: a Kernel PerspectiveAlberto BiettiICLR 2022 · 被引用 33 次
相关 Paper
- Role of Locality and Weight Sharing in Image-Based Tasks: A Sample Complexity Separation between CNNs, LCNs, and FCNsAakash Lahoti, Stefani Karp, Ezra Winston, Aarti Singh 等ICLR 2024 · 被引用 5 次
- Batch normalization is sufficient for universal function approximation in CNNsRebekka BurkholzICLR 2024 · 被引用 8 次
- Computational Separation Between Convolutional and Fully-Connected NetworksEran Malach, Shai Shalev-ShwartzICLR 2021 · 被引用 32 次
- Revisiting Spatial Invariance with Low-Rank Local ConnectivityGamaleldin F. Elsayed, Prajit Ramachandran, Jonathon Shlens, Simon KornblithICML 2020 · 被引用 51 次
- A unified theory of feature learning in RNNs and DNNsJan Bauer, Kirsten Fischer, Moritz Helias, Agostina PalmigianoICML 2026 · 被引用 4 次
