Lightweight Vision Transformer with Bidirectional Interaction
Qihang Fan, Huaibo Huang, Xiaoqiang Zhou, Ran He
摘要
Recent advancements in vision backbones have significantly improved their performance by simultaneously modeling images' local and global contexts. However, the bidirectional interaction between these two contexts has not been well explored and exploited, which is important in the human visual system. This paper proposes a Fully Adaptive Self-Attention (FASA) mechanism for vision transformer to model the local and global information as well as the bidirectional interaction between them in context-aware ways. Specifically, FASA employs self-modulated convolutions to adaptively extract local representation while utilizing self-attention in down-sampled space to extract global representation. Subsequently, it conducts a bidirectional adaptation process between local and global representation to model their interaction. In addition, we introduce a fine-grained downsampling strategy to enhance the down-sampled self-attention mechanism for finer-grained global perception capability. Based on FASA, we develop a family of lightweight vision backbones, Fully Adaptive Transformer (FAT) family. Extensive experiments on multiple vision tasks demonstrate that FAT achieves impressive performance. Notably, FAT accomplishes a 77.6% accuracy on ImageNet-1K using only 4.5M parameters and 0.7G FLOPs, which surpasses the most advanced ConvNets and Transformers with similar model size and computational costs. Moreover, our model exhibits faster speed on modern GPU compared to other models. Code will be available at https://github.com/qhfan/FAT .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Rectifying Magnitude Neglect in Linear AttentionQihang Fan, Huaibo Huang, Yuang Ai, Ran HeICCV 2025 · 被引用 14 次
- LookHere: Vision Transformers with Directed Attention Generalize and ExtrapolateAnthony Fuller, Daniel G. Kyrollos, Yousef Yassin, James R. GreenNeurIPS 2024 · 被引用 10 次
- MHLA: Restoring Expressivity of Linear Attention via Token-Level Multi-HeadKewei Zhang, Ye Huang, Yufan Deng, Jincheng Yu 等ICLR 2026 · 被引用 6 次
- Vision Transformer with Sparse Scan PriorYuguang Zhang, Qihang Fan, Huaibo HuangACM MM 2025 · 被引用 2 次
- Semantic Equitable Clustering: A Simple and Effective Strategy for Clustering Vision TokensQihang Fan, Huaibo Huang, Mingrui Chen, Ran HeICCV 2025 · 被引用 1 次
它引用的顶会 Paper48
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le 等ICCV 2019 · 被引用 9,163 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
相关 Paper
- Focal Attention for Long-Range Interactions in Vision TransformersJianwei Yang, Chunyuan Li, Pengchuan Zhang, Xiyang Dai 等NeurIPS 2021 · 被引用 228 次
- Lite Vision Transformer with Enhanced Self-AttentionChenglin Yang, Yilin Wang, Jianming Zhang, He Zhang 等CVPR 2022 · 被引用 139 次
- ViTAE: Vision Transformer Advanced by Exploring Intrinsic Inductive BiasYufei Xu, Qiming Zhang, Jing Zhang, Dacheng TaoNeurIPS 2021 · 被引用 429 次
- Pale Transformer: A General Vision Transformer Backbone with Pale-Shaped AttentionSitong Wu, Tianyi Wu, Haoru Tan, Guodong GuoAAAI 2022 · 被引用 84 次
- NomMer: Nominate Synergistic Context in Vision Transformer for Visual RecognitionHao Liu, Xinghua Jiang, Xin Li, Zhimin Bao 等CVPR 2022 · 被引用 17 次
