Non-deep Networks
Ankit Goyal, Alexey Bochkovskiy, Jia Deng, Vladlen Koltun
Abstract
Depth is the hallmark of deep neural networks. But more depth means more sequential computation and higher latency. This begs the question -is it possible to build high-performing "non-deep" neural networks? We show that it is. To do so, we use parallel subnetworks instead of stacking one layer after another. This helps effectively reduce depth while maintaining high performance. By utilizing parallel substructures, we show, for the first time, that a network with a depth of just 12 can achieve top-1 accuracy over 80% on ImageNet, 96% on CI-FAR10, and 81% on CIFAR100. We also show that a network with a low-depth ( 12 ) backbone can achieve an AP of 48% on MS-COCO. We analyze the scaling rules for our design and show how to increase performance without changing the network's depth. Finally, we provide a proof of concept for how non-deep networks could be used to build low-latency recognition systems. Code is available at https://github.com/imankgoyal/NonDeepNetworks .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 75de0bf7-55dc-444b-90ec-5c85eca0e2b9Cited by top-tier papers3
- Don't be lazy: CompleteP enables compute-efficient deep transformersNolan Dey, Bin Claire Zhang, Lorenzo Noci, Mufan Bill Li et al.NeurIPS 2025 · 77 citations
- Construction of Hierarchical Neural Architecture Search Spaces based on Context-free GrammarsSimon Schrodi, Danny Stoll, Binxin Ru, Rhea Sanjay Sukthanker et al.NeurIPS 2023 · 14 citations
- Channel-Spatial Support-Query Cross-Attention for Fine-Grained Few-Shot Image ClassificationShicheng Yang, Xiaoxu Li, Dongliang Chang, Zhanyu Ma et al.ACM MM 2024 · 12 citations
Builds on4
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- Scaled-YOLOv4: Scaling Cross Stage Partial NetworkChien-Yao Wang, Alexey Bochkovskiy, Hong-Yuan Mark LiaoCVPR 2021
- RepVGG: Making VGG-Style ConvNets Great AgainXiaohan Ding, Xiangyu Zhang, Ningning Ma, Jungong Han et al.CVPR 2021
Related papers
- Dynamic Convolution: Attention Over Convolution KernelsYinpeng Chen, Xiyang Dai, Mengchen Liu, Dongdong Chen et al.CVPR 2020
- Deep Isometric Learning for Visual RecognitionHaozhi Qi, Chong You, Xiaolong Wang, Yi Ma et al.ICML 2020 · 57 citations
- MobileOne: An Improved One millisecond Mobile BackbonePavan Kumar Anasosalu Vasu, James Gabriel, Jeff Zhu, Oncel Tuzel et al.CVPR 2023
- Efficient Latency-Aware CNN Depth Compression via Two-Stage Dynamic ProgrammingJinuk Kim, Yeonwoo Jeong, Deokjae Lee, Hyun Oh SongICML 2023 · 1 citation
- Is normalization indispensable for training deep neural network?Jie Shao, Kai Hu, Changhu Wang, Xiangyang Xue et al.NeurIPS 2020 · 70 citations
