SI-BiViT: Binarizing Vision Transformers with Spatial Interaction
Peng Yin, Xiaosu Zhu, Jingkuan Song, Lianli Gao, Heng Tao Shen
Abstract
Binarized Vision Transformers (BiViTs) aim to facilitate the efficient and lightweight utilization of Vision Transformers (ViTs) on devices with limited computational resources. Yet, the current approach to binarizing ViT leads to a substantial performance decrease compared to the full-precision model, posing obstacles to practical deployment. By empirical study, we reveal that spatial interaction (SI) is a critical factor that impacts performance due to lack of token-level correlation, but previous work ignores this factor. To this end, we design a ViT binarization approach dubbed SI-BiViT to incorporate spatial interaction in the binarization process. Specifically, an SI module is placed alongside the Multi-Layer Perceptron (MLP) module to formulate the dual-branch structure. This structure not only leverages knowledge from pre-trained ViTs by distilling over the original MLP, but also enhances spatial interaction via the introduced SI module. Correspondingly, we design a decoupled training strategy to train these two branches more effectively. Importantly, our SI-BiViT is orthogonal to existing Binarized ViTs approaches and can be directly plugged. Extensive experiments demonstrate the strong flexibility and effectiveness of SI-BiViT by plugging our method into four classic ViT backbones in supporting three downstream tasks, including classification, detection, and segmentation. In particular, SI-BiViT enhances the classification performance of binarized ViTs by an average of 10.52% in Top-1 accuracy compared to the previous state-of-the-art. Codes are available at https://github.com/VL-Group/SI-BiViT
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 230d9954-2d95-49ed-8a4e-ca213c012d91Cited by top-tier papers2
- Information-Bottleneck Driven Binary Neural Network for Change DetectionKaijie Yin, Zhiyuan Zhang, Shu Kong, Tian Gao et al.ICCV 2025 · 4 citations
- BHViT: Binarized Hybrid Vision TransformerTian Gao, Yu Zhang, Zhiyuan Zhang, Huajun Liu et al.CVPR 2025
Related papers
- BiViT: Extremely Compressed Binary Vision TransformersYefei He, Zhenyu Lou, Luoming Zhang, Jing Liu et al.ICCV 2023 · 44 citations
- EViT: Expediting Vision Transformers via Token ReorganizationsYouwei Liang, Chongjian Ge, Zhan Tong, Yibing Song et al.ICLR 2022 · 137 citations
- MG-ViT: A Multi-Granularity Method for Compact and Efficient Vision TransformersYu Zhang, Yepeng Liu, Duoqian Miao, Qi Zhang et al.NeurIPS 2023 · 23 citations
- Bi-ViT: Pushing the Limit of Vision Transformer QuantizationYanjing Li, Sheng Xu, Mingbao Lin, Xianbin Cao et al.AAAI 2024 · 23 citations
- Instance-Aware Group Quantization for Vision TransformersJaehyeon Moon, Dohyung Kim, Junyong Cheon, Bumsub HamCVPR 2024 · 11 citations
