Efficient Deformable ConvNets: Rethinking Dynamic and Sparse Operator for Vision Applications
Yuwen Xiong, Zhiqi Li, Yuntao Chen, Feng Wang, Xizhou Zhu, Jiapeng Luo, Wenhai Wang, Tong Lu, Hongsheng Li, Yu Qiao, Lewei Lu, Jie Zhou, Jifeng Dai
Abstract
We introduce Deformable Convolution v4 (DCNv4), a highly efficient and effective operator designed for a broad spectrum of vision applications. DCNv4 addresses the limitations of its predecessor, DCNv3, with two key enhancements: 1. removing softmax normalization in spatial aggregation to enhance its dynamic property and expressive power and 2. optimizing memory access to minimize redundant operations for speedup. These improvements result in a significantly faster convergence compared to DCNv3 and a substantial increase in processing speed, with DCNv4 achieving more than three times the forward speed. DCNv4 demonstrates exceptional performance across various tasks, including image classification, instance and semantic segmentation, and notably, image generation. When integrated into generative models like U-Net in the latent diffusion model, DCNv4 outperforms its baseline, underscoring its possibility to enhance generative models. In practical applications, replacing DCNv3 with DCNv4 in the InternImage model to create FlashInternImage results in up to 80% speed increase and further performance improvement without further modifications. The advancements in speed and efficiency of DCNv4, combined with its robust performance across diverse vision tasks, show its potential as a foundational building block for future vision models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 53f572df-5f8b-4b1c-afd5-7b8215bac394Cited by top-tier papers23
- Exploring DCN-like architecture for fast image generation with arbitrary resolutionShuai Wang, Zexian Li, Tianhui Song, Xubin Li et al.NeurIPS 2024 · 16 citations
- Vision-RWKV: Efficient and Scalable Visual Perception with RWKV-Like ArchitecturesYuchen Duan, Weiyun Wang, Zhe Chen, Xizhou Zhu et al.ICLR 2025 · 10 citations
- Diff2I2P: Differentiable Image-to-Point Cloud Registration with Diffusion PriorJuncheng Mu, Chengwei Ren, Weixiang Zhang, Liang Pan et al.ICCV 2025 · 9 citations
- UniDAC: Universal Metric Depth Estimation for Any CameraGirish Chandar Ganesan, Yuliang Guo, Liu Ren, Xiaoming LiuCVPR 2026 · 8 citations
- Residual Diffusion Deblurring Model for Single Image Defocus DeblurringHaoxuan Feng, Haohui Zhou, Tian Ye, Sixiang Chen et al.AAAI 2025 · 5 citations
Builds on18
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
Related papers
- DefT: Boosting Scalability of Deformable Convolution Operations on GPUsEdward Hanson, Mark Horton, Hai (Helen) Li, Yiran ChenASPLOS 2023 · 1 citation
- Run, Don't Walk: Chasing Higher FLOPS for Faster Neural NetworksJierun Chen, Shiu-Hong Kao, Hao He, Weipeng Zhuo et al.CVPR 2023
- InternImage: Exploring Large-Scale Vision Foundation Models with Deformable ConvolutionsWenhai Wang, Jifeng Dai, Zhe Chen, Zhenhang Huang et al.CVPR 2023
- DiC: Rethinking Conv3x3 Designs in Diffusion ModelsYuchuan Tian, Jing Han, Chengcheng Wang, Yuchen Liang et al.CVPR 2025
- ELFATT: Efficient Linear Fast Attention for Vision TransformersChong Wu, Maolin Che, Renjie Xu, Zhuoheng Ran et al.ACM MM 2025 · 3 citations
