DefT: Boosting Scalability of Deformable Convolution Operations on GPUs
Edward Hanson, Mark Horton, Hai (Helen) Li, Yiran Chen
摘要
Deformable Convolutional Networks (DCN) have been proposed as a powerful tool to boost the representation power of Convolutional Neural Networks (CNN) in computer vision tasks via adaptive sampling of the input feature map. Much like vision transformers, DCNs utilize a more flexible inductive bias than standard CNNs and have also been shown to improve performance of particular models. For example, drop-in DCN layers were shown to increase the AP score of Mask RCNN by 10.6 points while introducing only 1% additional parameters and FLOPs, improving the state-of-the-art model at the time of publication. However, despite evidence that more DCN layers placed earlier in the network can further improve performance, we have not seen this trend continue with further scaling of deformations in CNNs, unlike for vision transformers. Benchmarking experiments show that a realistically sized DCN layer (64H×64W, 64 in-out channel) incurs a 4× slowdown on a GPU platform, discouraging the more ubiquitous use of deformations in CNNs. These slowdowns are caused by the irregular input-dependent access patterns of the bilinear interpolation operator, which has a disproportionately low arithmetic intensity (AI) compared to the rest of the DCN. To address the disproportionate slowdown of DCNs and enable their expanded use in CNNs, we propose DefT, a series of workload-aware optimizations for DCN kernels. DefT identifies performance bottlenecks in DCNs and fuses specific operators that are observed to limit DCN AI. Our approach also uses statistical information of DCN workloads to adapt the workload tiling to the DCN layer dimensions, minimizing costly out-of-boundary input accesses. Experimental results show that DefT mitigates up to half of DCN slowdown over the current-art PyTorch implementation. This translates to a layerwise speedup of up to 134% and a reduction of normalized training time of 46% on a fully DCN-enabled ResNet model.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Fast Fourier ConvolutionLu Chi, Borui Jiang, Yadong MuNeurIPS 2020 · 被引用 842 次
- Vision Transformer with Deformable AttentionZhuofan Xia, Xuran Pan, Shiji Song, Li Erran Li 等CVPR 2022 · 被引用 835 次
- DNNFusion: accelerating deep neural networks execution with advanced operator fusionWei Niu, Jiexiong Guan, Yanzhi Wang, Gagan Agrawal 等PLDI 2021 · 被引用 166 次
- GCNAX: A Flexible and Energy-efficient Accelerator for Graph Convolutional Neural NetworksJiajun Li, Ahmed Louri, Avinash Karanth, Razvan C. BunescuHPCA 2021 · 被引用 147 次
相关 Paper
- DeMT: Deformable Mixer Transformer for Multi-Task Learning of Dense PredictionYangyang Xu, Yibo Yang, Lefei ZhangAAAI 2023 · 被引用 81 次
- DEFA: Efficient Deformable Attention Acceleration via Pruning-Assisted Grid-Sampling and Multi-Scale Parallel ProcessingYansong Xu, Dongxu Lyu, Zhenyu Li, Yuzhou Chen 等DAC 2024 · 被引用 5 次
- Deformable Kernels: Adapting Effective Receptive Fields for Object DeformationHang Gao, Xizhou Zhu, Stephen Lin, Jifeng DaiICLR 2020 · 被引用 72 次
- Efficient Deformable ConvNets: Rethinking Dynamic and Sparse Operator for Vision ApplicationsYuwen Xiong, Zhiqi Li, Yuntao Chen, Feng Wang 等CVPR 2024 · 被引用 205 次
- InternImage: Exploring Large-Scale Vision Foundation Models with Deformable ConvolutionsWenhai Wang, Jifeng Dai, Zhe Chen, Zhenhang Huang 等CVPR 2023
