UniNeXt: Exploring A Unified Architecture for Vision Recognition
Fangjian Lin, Jianlong Yuan, Sitong Wu, Fan Wang, Zhibin Wang
Abstract
Vision Transformers have shown great potential in computer vision tasks. Most recent works have focused on elaborating the spatial token mixer for performance gains. However, we observe that a well-designed general architecture can significantly improve the performance of the entire backbone, regardless of which spatial token mixer is equipped. In this paper, we propose UniNeXt, an improved general architecture for the vision backbone. To verify its effectiveness, we instantiate the spatial token mixer with various typical and modern designs, including both convolution and attention modules. Compared with the architecture in which they are first proposed, our UniNeXt architecture can steadily boost the performance of all the spatial token mixers, and narrows the performance gap among them. Surprisingly, our UniNeXt equipped with naive local window attention even outperforms the previous state-of-the-art. Interestingly, the ranking of these spatial token mixers also changes under our UniNeXt, suggesting that an excellent spatial token mixer may be stifled due to a suboptimal general architecture, which further shows the importance of the study on the general architecture of vision backbone. Code is available at UniNeXt.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- CogVLM: Visual Expert for Pretrained Language ModelsWeihan Wang, Qingsong Lv, Wenmeng Yu, Wenyi Hong et al.NeurIPS 2024 · 858 citations
- Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMsYongyi Su, Haojie Zhang, Shijie Li, Nanqing Liu et al.ICLR 2026 · 22 citations
- Unleashing the Potential of Multimodal LLMs for Zero-Shot Spatio-Temporal Video GroundingZaiquan Yang, Yuhao Liu, Gerhard P. Hancke, Rynson W. H. LauNeurIPS 2025 · 10 citations
- SaCo Loss: Sample-Wise Affinity Consistency for Vision-Language Pre-TrainingSitong Wu, Haoru Tan, Zhuotao Tian, Yukang Chen et al.CVPR 2024 · 5 citations
- TMT-VIS: Taxonomy-aware Multi-dataset Joint Training for Video Instance SegmentationRongkun Zheng, Lu Qi, Xi Chen, Yi Wang et al.NeurIPS 2023 · 3 citations
Builds on13
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- Twins: Revisiting the Design of Spatial Attention in Vision TransformersXiangxiang Chu, Zhi Tian, Yuqing Wang, Bo Zhang et al.NeurIPS 2021 · 1,388 citations
- CSWin Transformer: A General Vision Transformer Backbone with Cross-Shaped WindowsXiaoyi Dong, Jianmin Bao, Dongdong Chen, Weiming Zhang et al.CVPR 2022 · 1,207 citations
Related papers
- MetaFormer is Actually What You Need for VisionWeihao Yu, Mi Luo, Pan Zhou, Chenyang Si et al.CVPR 2022 · 1,114 citations
- ResT: An Efficient Transformer for Visual RecognitionQinglong Zhang, Yu-Bin YangNeurIPS 2021 · 313 citations
- Active Token MixerGuoqiang Wei, Zhizheng Zhang, Cuiling Lan, Yan Lu et al.AAAI 2023 · 25 citations
- RIFormer: Keep Your Vision Backbone Effective But Removing Token MixerJiahao Wang, Songyang Zhang, Yong Liu, Taiqiang Wu et al.CVPR 2023
- Rethinking Spatial Dimensions of Vision TransformersByeongho Heo, Sangdoo Yun, Dongyoon Han, Sanghyuk Chun et al.ICCV 2021 · 733 citations
