ViT-CoMer: Vision Transformer with Convolutional Multi-scale Feature Interaction for Dense Predictions
Chunlong Xia, Xinliang Wang, Feng Lv, Xin Hao, Yifeng Shi
Abstract
Although Vision Transformer (ViT) has achieved significant success in computer vision, it does not perform well in dense prediction tasks due to the lack of inner-patch information interaction and the limited diversity of feature scale. Most existing studies are devoted to designing vision-specific transformers to solve the above problems, which introduce additional pre-training costs. Therefore, we present a plain, pre-training-free, and feature-enhanced ViT back-bone with Convolutional Multi-scale feature interaction, named ViT-CoMer, which facilitates bidirectional interaction between CNN and transformer. Compared to the state-of-the-art, ViT-CoMer has the following advantages: (1) We inject spatial pyramid multi-receptive field convolutional features into the ViT architecture, which effectively alleviates the problems of limited local information interaction and single-feature representation in ViT. (2) We propose a simple and efficient CNN- Transformer bidirectional fusion interaction module that performs multi-scale fusion across hierarchical features, which is beneficial for han-dling dense prediction tasks. (3) We evaluate the performance of ViT-CoMer across various dense prediction tasks, different frameworks, and multiple advanced pre-training. Notably, our ViT-CoMer-L achieves 64.3% AP on COCO val2017 without extra training data, and 62.1% mIoU on ADE20K val, both of which are comparable to state-of-the-art methods. We hope ViT-CoMer can serve as a new backbone for dense prediction tasks to facilitate future research. The code will be released at https://github.com/Traffic-xlviT-CoMer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 544c314d-afe9-4b9a-b902-fd39506b096bCited by top-tier papers23
- UniRGB-IR: A Unified Framework for Visible-Infrared Semantic Tasks via Adapter TuningMaoxun Yuan, Bo Cui, Tianyi Zhao, Jiayi Wang et al.ACM MM 2025 · 21 citations
- Depth Pro: Sharp Monocular Metric Depth in Less Than a SecondAlexey Bochkovskiy, Amaël Delaunoy, Hugo Germain, Marcel Santos et al.ICLR 2025 · 15 citations
- Rethinking Multi-Modal Object Detection From the Perspective of Mono-Modality Feature LearningTianyi Zhao, Boyang Liu, Yanglei Gao, Yiming Sun et al.ICCV 2025 · 15 citations
- Frequency-Dynamic Attention Modulation for Dense PredictionLinwei Chen, Lin Gu, Ying FuICCV 2025 · 13 citations
- Parameter-Inverted Image Pyramid NetworksXizhou Zhu, Xue Yang, Zhaokai Wang, Hao Li et al.NeurIPS 2024 · 11 citations
Builds on26
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
Related papers
- Vision Transformer Adapter for Dense PredictionsZhe Chen, Yuchen Duan, Wenhai Wang, Junjun He et al.ICLR 2023 · 204 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- MPViT: Multi-Path Vision Transformer for Dense PredictionYoungwan Lee, Jonghee Kim, Jeffrey Willette, Sung Ju HwangCVPR 2022 · 339 citations
- The Missing Point in Vision Transformers for Universal Image SegmentationSajjad Shahabodini, Mobina Mansoori, Farnoush Bayatmakou, Jamshid Abouei et al.CVPR 2026 · 5 citations
- Global Context Vision TransformersAli Hatamizadeh, Hongxu Yin, Greg Heinrich, Jan Kautz et al.ICML 2023 · 213 citations
