TransNeXt: Robust Foveal Visual Perception for Vision Transformers
Dai Shi
Abstract
Due to the depth degradation effect in residual connections, many efficient Vision Transformers models that rely on stacking layers for information exchange often fail to form sufficient information mixing, leading to unnatural visual perception. To address this issue, in this paper, we propose Aggregated Attention, a biomimetic design-based token mixer that simulates biological foveal vision and continuous eye movement while enabling each token on the feature map to have a global perception. Furthermore, we incorporate learnable tokens that interact with conventional queries and keys, which further diversifies the generation of affinity matrices beyond merely relying on the similarity between queries and keys. Our approach does not rely on stacking for information exchange, thus effectively avoiding depth degradation and achieving natural visual perception. Additionally, we propose Convolutional GLU, a channel mixer that bridges the gap between GLU and SE mechanism, which empowers each token to have channel attention based on its nearest neighbor image features, enhancing local modeling capability and model robustness. We combine aggregated attention and convolutional GLU to create a new visual backbone called TransNeXt. Extensive experiments demonstrate that our TransNeXt achieves state-of-the-art performance across multiple model sizes. At a resolution of 224<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">2</sup>, TransNeXt-Tiny attains an ImageNet accuracy of 84.0%, surpassing ConvNeXt-B with 69% fewer parameters. Our TransNeXt-Base achieves an ImageNet accuracy of 86.2% and an ImageNet-A accuracy of 61. 6% at a resolution of 384<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">2</sup>, a COCO object detection mAP of 57.1, and an ADE20K semantic segmentation mIoU of 54. 7.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0e5f3abe-1cb5-46f5-95de-fa62e621377bCited by top-tier papers23
- Mamba YOLO: A Simple Baseline for Object Detection with State Space ModelZeyu Wang, Chen Li, Huiying Xu, Xinzhong Zhu et al.AAAI 2025 · 136 citations
- Unleashing the Power of Generic Segmentation Model: A Simple Baseline for Infrared Small Target DetectionMingjin Zhang, Chi Zhang, Qiming Zhang, Yunsong Li et al.ACM MM 2024 · 33 citations
- DAMamba: Vision State Space Model with Dynamic Adaptive ScanTanzhe Li, Caoshuo Li, Jiayi Lyu, Hongjuan Pei et al.NeurIPS 2025 · 24 citations
- VSSD: Vision Mamba With Non-Causal State Space DualityYuheng Shi, Mingjia Li, Minjing Dong, Chang XuICCV 2025 · 20 citations
- Cross Paradigm Representation and Alignment Transformer for Image DerainingShun Zou, Yi Zou, Juncheng Li, Guangwei Gao et al.ACM MM 2025 · 19 citations
Builds on27
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
Related papers
- UniNeXt: Exploring A Unified Architecture for Vision RecognitionFangjian Lin, Jianlong Yuan, Sitong Wu, Fan Wang et al.ACM MM 2023 · 15 citations
- SegNeXt: Rethinking Convolutional Attention Design for Semantic SegmentationMeng-Hao Guo, Cheng-Ze Lu, Qibin Hou, Zhengning Liu et al.NeurIPS 2022 · 1,385 citations
- Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNetLi Yuan, Yunpeng Chen, Tao Wang, Weihao Yu et al.ICCV 2021 · 2,462 citations
- FastViT: A Fast Hybrid Vision Transformer using Structural ReparameterizationPavan Kumar Anasosalu Vasu, James Gabriel, Jeff Zhu, Oncel Tuzel et al.ICCV 2023 · 341 citations
- OverLoCK: An Overview-first-Look-Closely-next ConvNet with Context-Mixing Dynamic KernelsMeng Lou, Yizhou YuCVPR 2025
