SparseFormer: Sparse Visual Recognition via Limited Latent Tokens
Ziteng Gao, Zhan Tong, Limin Wang, Mike Zheng Shou
摘要
Human visual recognition is a sparse process, where only a few salient visual cues are attended to rather than traversing every detail uniformly. However, most current vision networks follow a dense paradigm, processing every single visual unit (e.g., pixel or patch) in a uniform manner. In this paper, we challenge this dense paradigm and present a new method, coined SparseFormer, to imitate human's sparse visual recognition in an end-to-end manner. Sparse-Former learns to represent images using a highly limited number of tokens (down to 49) in the latent space with sparse feature sampling procedure instead of processing dense units in the original pixel space. Therefore, Sparse-Former circumvents most of dense operations on the image space and has much lower computational costs. Experiments on the ImageNet classification benchmark dataset show that SparseFormer achieves performance on par with canonical or well-established models while offering better accuracy-throughput tradeoff. Moreover, the design of our network can be easily extended to the video classification with promising performance at lower computational costs. We hope that our work can provide an alternative way for visual modeling and inspire further research on sparse neural architectures. The code will be publicly available at https://github.com/showlab/sparseformer .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Matryoshka Query Transformer for Large Vision-Language ModelsWenbo Hu, Zi-Yi Dou, Liunian Harold Li, Amita Kamath 等NeurIPS 2024 · 被引用 58 次
- Revisiting Vision Transformer from the View of Path EnsembleShuning Chang, Pichao Wang, Hao Luo, Fan Wang 等ICCV 2023 · 被引用 8 次
- Sparse Image Synthesis via Joint Latent and RoI FlowZiteng Gao, Jay Zhangjie Wu, Mike Zheng ShouNeurIPS 2025
- Less Is Better: Sparse Instance Learning for Cross-Domain Few-Shot Object DetectionYali Huang, Jie Mei, Ziyi Wu, Yiming Yang 等AAAI 2026
- Bootstrapping SparseFormers from Vision Foundation ModelsZiteng Gao, Zhan Tong, Kevin Qinghong Lin, Joya Chen 等CVPR 2024
它引用的顶会 Paper23
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
相关 Paper
- Scalable Vision Transformers with Hierarchical PoolingZizheng Pan, Bohan Zhuang, Jing Liu, Haoyu He 等ICCV 2021 · 被引用 154 次
- MetaFormer is Actually What You Need for VisionWeihao Yu, Mi Luo, Pan Zhou, Chenyang Si 等CVPR 2022 · 被引用 1,114 次
- Spikformer: When Spiking Neural Network Meets TransformerZhaokun Zhou, Yuesheng Zhu, Chao He, Yaowei Wang 等ICLR 2023 · 被引用 103 次
- Associative Memory Augmented Asynchronous Spatiotemporal Representation Learning for Event-based PerceptionUday Kamal, Saurabh Dash, Saibal MukhopadhyayICLR 2023
- Vision Transformer with Sparse Scan PriorYuguang Zhang, Qihang Fan, Huaibo HuangACM MM 2025 · 被引用 2 次
