CF-ViT: A General Coarse-to-Fine Method for Vision Transformer
Mengzhao Chen, Mingbao Lin, Ke Li, Yunhang Shen, Yongjian Wu, Fei Chao, Rongrong Ji
摘要
Vision Transformers (ViT) have made many breakthroughs in computer vision tasks. However, considerable redundancy arises in the spatial dimension of an input image, leading to massive computational costs. Therefore, We propose a coarse-to-fine vision transformer (CF-ViT) to relieve computational burden while retaining performance in this paper. Our proposed CF-ViT is motivated by two important observations in modern ViT models: (1) The coarse-grained patch splitting can locate informative regions of an input image. (2) Most images can be well recognized by a ViT model in a small-length token sequence. Therefore, our CF-ViT implements network inference in a two-stage manner. At coarse inference stage, an input image is split into a small-length patch sequence for a computationally economical classification. If not well recognized, the informative patches are identified and further re-split in a fine-grained granularity. Extensive experiments demonstrate the efficacy of our CF-ViT. For example, without any compromise on performance, CF-ViT reduces 53% FLOPs of LV-ViT, and also achieves 2.01× throughput. Code of this project is at https://github.com/ChenMnZ/CF-ViT .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper26
- MogaNet: Multi-order Gated Aggregation NetworkSiyuan Li, Zedong Wang, Zicheng Liu, Cheng Tan 等ICLR 2024 · 被引用 151 次
- DiffRate : Differentiable Compression Rate for Efficient Vision TransformersMengzhao Chen, Wenqi Shao, Peng Xu, Mingbao Lin 等ICCV 2023 · 被引用 87 次
- SHaRPose: Sparse High-Resolution Representation for Human Pose EstimationXiaoqi An, Lin Zhao, Chen Gong, Nannan Wang 等AAAI 2024 · 被引用 36 次
- Leveraging Vision-Centric Multi-Modal Expertise for 3D Object DetectionLinyan Huang, Zhiqi Li, Chonghao Sima, Wenhai Wang 等NeurIPS 2023 · 被引用 26 次
- LF-ViT: Reducing Spatial Redundancy in Vision Transformer for Efficient Image RecognitionYoubing Hu, Yun Cheng, Anqi Lu, Zhiqiang Cao 等AAAI 2024 · 被引用 25 次
它引用的顶会 Paper25
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image ClassificationChun-Fu (Richard) Chen, Quanfu Fan, Rameswar PandaICCV 2021 · 被引用 2,072 次
- DynamicViT: Efficient Vision Transformers with Dynamic Token SparsificationYongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu 等NeurIPS 2021 · 被引用 1,343 次
- Going deeper with Image TransformersHugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve 等ICCV 2021 · 被引用 1,279 次
相关 Paper
- MG-ViT: A Multi-Granularity Method for Compact and Efficient Vision TransformersYu Zhang, Yepeng Liu, Duoqian Miao, Qi Zhang 等NeurIPS 2023 · 被引用 23 次
- Scalable Vision Transformers with Hierarchical PoolingZizheng Pan, Bohan Zhuang, Jing Liu, Haoyu He 等ICCV 2021 · 被引用 154 次
- EViT: Expediting Vision Transformers via Token ReorganizationsYouwei Liang, Chongjian Ge, Zhan Tong, Yibing Song 等ICLR 2022 · 被引用 137 次
- Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNetLi Yuan, Yunpeng Chen, Tao Wang, Weihao Yu 等ICCV 2021 · 被引用 2,462 次
- A-ViT: Adaptive Tokens for Efficient Vision TransformerHongxu Yin, Arash Vahdat, José M. Álvarez, Arun Mallya 等CVPR 2022 · 被引用 288 次
