Learning Strides in Convolutional Neural Networks
Rachid Riad, Olivier Teboul, David Grangier, Neil Zeghidour
摘要
Convolutional neural networks typically contain several downsampling operators, such as strided convolutions or pooling layers, that progressively reduce the resolution of intermediate representations. This provides some shift-invariance while reducing the computational complexity of the whole architecture. A critical hyperparameter of such layers is their stride: the integer factor of downsampling. As strides are not differentiable, finding the best configuration either requires cross-validation or discrete optimization (e.g. architecture search), which rapidly become prohibitive as the search space grows exponentially with the number of downsampling layers. Hence, exploring this search space by gradient descent would allow finding better configurations at a lower computational cost. This work introduces DiffStride, the first downsampling layer with learnable strides. Our layer learns the size of a cropping mask in the Fourier domain, that effectively performs resizing in a differentiable way. Experiments on audio and image classification show the generality and effectiveness of our solution: we use DiffStride as a drop-in replacement to standard downsampling layers and outperform them. In particular, we show that introducing our layer into a ResNet-18 architecture allows keeping consistent high performance on CIFAR10, CIFAR100 and ImageNet even when training starts from poor random stride configurations. Moreover, formulating strides as learnable variables allows us to introduce a regularization term that controls the computational complexity of the architecture. We show how this regularization allows trading off accuracy for efficiency on ImageNet.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- PointTAD: Multi-Label Temporal Action Detection with Learnable Query PointsJing Tan, Xiaotong Zhao, Xintian Shi, Bin Kang 等NeurIPS 2022 · 被引用 41 次
- Dilated convolution with learnable spacingsIsmail Khalfaoui Hassani, Thomas Pellegrini, Timothée MasquelierICLR 2023 · 被引用 18 次
- Learning Temporal Resolution in Spectrogram for Audio ClassificationHaohe Liu, Xubo Liu, Qiuqiang Kong, Wenwu Wang 等AAAI 2024 · 被引用 15 次
- Pooling Revisited: Your Receptive Field is SuboptimalDong-Hwan Jang, Sanghyeok Chu, Joonhyuk Kim, Bohyung HanCVPR 2022 · 被引用 13 次
- FouriDown: Factoring Down-Sampling into Shuffling and SuperposingQi Zhu, Man Zhou, Jie Huang, Naishan Zheng 等NeurIPS 2023 · 被引用 12 次
它引用的顶会 Paper8
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le 等ICCV 2019 · 被引用 9,163 次
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 被引用 1,861 次
- CoAtNet: Marrying Convolution and Attention for All Data SizesZihang Dai, Hanxiao Liu, Quoc V. Le, Mingxing TanNeurIPS 2021 · 被引用 1,747 次
- Revisiting ResNets: Improved Training and Scaling StrategiesIrwan Bello, William Fedus, Xianzhi Du, Ekin Dogus Cubuk 等NeurIPS 2021 · 被引用 378 次
相关 Paper
- Learning in the Frequency DomainKai Xu, Minghai Qin, Fei Sun, Yuhao Wang 等CVPR 2020
- Truly Scale-Equivariant Deep Nets with Fourier LayersMd Ashiqur Rahman, Raymond A. YehNeurIPS 2023 · 被引用 17 次
- FlexConv: Continuous Kernel Convolutions With Differentiable Kernel SizesDavid W. Romero, Robert-Jan Bruintjes, Jakub Mikolaj Tomczak, Erik J. Bekkers 等ICLR 2022 · 被引用 94 次
- Efficient Architecture Search for Diverse TasksJunhong Shen, Mikhail Khodak, Ameet TalwalkarNeurIPS 2022 · 被引用 42 次
- ShiftMorph: A Fast and Robust Convolutional Neural Network for 3D Deformable Medical Image RegistrationLijian Yang, Weisheng Li, Yucheng Shu, Jian-Xun Mi 等ACM MM 2024 · 被引用 4 次
