ResT V2: Simpler, Faster and Stronger
Qinglong Zhang, Yu-Bin Yang
摘要
This paper proposes ResTv2, a simpler, faster, and stronger multi-scale vision Transformer for visual recognition. ResTv2 simplifies the EMSA structure in ResTv1 (i.e., eliminating the multi-head interaction part) and employs an upsample operation to reconstruct the lost medium- and high-frequency information caused by the downsampling operation. In addition, we explore different techniques for better apply ResTv2 backbones to downstream tasks. We found that although combining EMSAv2 and window attention can greatly reduce the theoretical matrix multiply FLOPs, it may significantly decrease the computation density, thus causing lower actual speed. We comprehensively validate ResTv2 on ImageNet classification, COCO detection, and ADE20K semantic segmentation. Experimental results show that the proposed ResTv2 can outperform the recently state-of-the-art backbones by a large margin, demonstrating the potential of ResTv2 as solid backbones. The code and models will be made publicly available at https://github.com/wofmanaf/ResT
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- SG-Former: Self-guided Transformer with Evolving Token ReallocationSucheng Ren, Xingyi Yang, Songhua Liu, Xinchao WangICCV 2023 · 被引用 70 次
- Accelerating Pre-training of Multimodal LLMs via Chain-of-SightZiyuan Huang, Kaixiang Ji, Biao Gong, Zhiwu Qing 等NeurIPS 2024 · 被引用 9 次
- Intern VL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic TasksZhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su 等CVPR 2024
它引用的顶会 Paper30
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
相关 Paper
- ResT: An Efficient Transformer for Visual RecognitionQinglong Zhang, Yu-Bin YangNeurIPS 2021 · 被引用 313 次
- MViTv2: Improved Multiscale Vision Transformers for Classification and DetectionYanghao Li, Chao-Yuan Wu, Haoqi Fan, Karttikeya Mangalam 等CVPR 2022 · 被引用 699 次
- MPViT: Multi-Path Vision Transformer for Dense PredictionYoungwan Lee, Jonghee Kim, Jeffrey Willette, Sung Ju HwangCVPR 2022 · 被引用 339 次
- ViT-CoMer: Vision Transformer with Convolutional Multi-scale Feature Interaction for Dense PredictionsChunlong Xia, Xinliang Wang, Feng Lv, Xin Hao 等CVPR 2024 · 被引用 109 次
- Scale-Aware Modulation Meet TransformerWeifeng Lin, Ziheng Wu, Jiayu Chen, Jun Huang 等ICCV 2023 · 被引用 154 次
