BVT-IMA: Binary Vision Transformer with Information-Modified Attention
Zhenyu Wang, Hao Luo, Xuemei Xie, Fan Wang, Guangming Shi
Abstract
As a compression method that can significantly reduce the cost of calculations and memories, model binarization has been extensively studied in convolutional neural networks. However, the recently popular vision transformer models pose new challenges to such a technique, in which the binarized models suffer from serious performance drops. In this paper, an attention shifting is observed in the binary multi-head self-attention module, which can influence the information fusion between tokens and thus hurts the model performance. From the perspective of information theory, we find a correlation between attention scores and the information quantity, further indicating that a reason for such a phenomenon may be the loss of the information quantity induced by constant moduli of binarized tokens. Finally, we reveal the information quantity hidden in the attention maps of binary vision transformers and propose a simple approach to modify the attention values with look-up information tables so that improve the model performance. Extensive experiments on CIFAR-100/TinyImageNet/ImageNet-1k demonstrate the effectiveness of the proposed information-modified attention on binary vision transformers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cfafc7d9-497c-420c-9d90-9b2d188395a7Cited by top-tier papers1
Ask how each one uses itBuilds on21
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Chasing Sparsity in Vision Transformers: An End-to-End ExplorationTianlong Chen, Yu Cheng, Zhe Gan, Lu Yuan et al.NeurIPS 2021 · 295 citations
- TokenLearner: Adaptive Space-Time Tokenization for VideosMichael S. Ryoo, A. J. Piergiovanni, Anurag Arnab, Mostafa Dehghani et al.NeurIPS 2021 · 274 citations
Related papers
- BiViT: Extremely Compressed Binary Vision TransformersYefei He, Zhenyu Lou, Luoming Zhang, Jing Liu et al.ICCV 2023 · 44 citations
- BinaryAttention: One-Bit QK-Attention for Vision and Diffusion TransformersChaodong XIAO, Zhengqiang ZHANG, Lei ZhangCVPR 2026 · 1 citation
- Frequency-Aware Token Reduction for Efficient Vision TransformerDongJae Lee, Jiwan Hur, Jaehyun Choi, Jaemyung Yu et al.NeurIPS 2025 · 4 citations
- Q-ViT: Accurate and Fully Quantized Low-bit Vision TransformerYanjing Li, Sheng Xu, Baochang Zhang, Xianbin Cao et al.NeurIPS 2022 · 185 citations
- Post-Training Quantization for Vision TransformerZhenhua Liu, Yunhe Wang, Kai Han, Wei Zhang et al.NeurIPS 2021 · 528 citations
