BiViT: Extremely Compressed Binary Vision Transformers
Yefei He, Zhenyu Lou, Luoming Zhang, Jing Liu, Weijia Wu, Hong Zhou, Bohan Zhuang
摘要
Model binarization can significantly compress model size, reduce energy consumption, and accelerate inference through efficient bit-wise operations. Although binarizing convolutional neural networks have been extensively studied, there is little work on exploring binarization of vision Transformers which underpin most recent breakthroughs in visual recognition. To this end, we propose to solve two fundamental challenges to push the horizon of Binary Vision Transformers (BiViT). First, the traditional binary method does not take the long-tailed distribution of softmax attention into consideration, bringing large binarization errors in the attention module. To solve this, we propose Softmax-aware Binarization, which dynamically adapts to the data distribution and reduces the error caused by binarization. Second, to better preserve the information of the pretrained model and restore accuracy, we propose a Cross-layer Binarization scheme that decouples the binarization of self-attention and multi-layer perceptrons (MLPs), and Parameterized Weight Scales which introduce learnable scaling factors for weight binarization. Overall, our method performs favorably against state-of-the-arts by 19.8% on the TinyImageNet dataset. On ImageNet, our BiViT achieves a competitive 75.6% Top-1 accuracy over Swin-S model. Additionally, on COCO object detection, our method achieves an mAP of 40.8 with a Swin-T backbone over Cascade Mask R-CNN framework.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Outlier-aware Slicing for Post-Training Quantization in Vision TransformerYuexiao Ma, Huixia Li, Xiawu Zheng, Feng Ling 等ICML 2024 · 被引用 17 次
- BiDM: Pushing the Limit of Quantization for Diffusion ModelsXingyu Zheng, Xianglong Liu, Yichen Bian, Xudong Ma 等NeurIPS 2024 · 被引用 12 次
- Training Binary Neural Networks via Gaussian Variational Inference and Low-Rank Semidefinite ProgrammingLorenzo Orecchia, Jiawei Hu, Xue He, Wang Mark 等NeurIPS 2024 · 被引用 4 次
- Efficiency Follows Global-Local DecouplingZhenyu Yang, Gensheng Pei, Tao Chen, Yichao Zhou 等CVPR 2026 · 被引用 3 次
- BinaryAttention: One-Bit QK-Attention for Vision and Diffusion TransformersChaodong XIAO, Zhengqiang ZHANG, Lei ZhangCVPR 2026 · 被引用 1 次
它引用的顶会 Paper24
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan 等ICCV 2021 · 被引用 4,909 次
- CoAtNet: Marrying Convolution and Attention for All Data SizesZihang Dai, Hanxiao Liu, Quoc V. Le, Mingxing TanNeurIPS 2021 · 被引用 1,747 次
相关 Paper
- Bi-ViT: Pushing the Limit of Vision Transformer QuantizationYanjing Li, Sheng Xu, Mingbao Lin, Xianbin Cao 等AAAI 2024 · 被引用 23 次
- MiniViT: Compressing Vision Transformers with Weight MultiplexingJinnian Zhang, Houwen Peng, Kan Wu, Mengchen Liu 等CVPR 2022 · 被引用 115 次
- Quantized Feature Distillation for Network QuantizationKe Zhu, Yin-Yin He, Jianxin WuAAAI 2023 · 被引用 21 次
- BHViT: Binarized Hybrid Vision TransformerTian Gao, Yu Zhang, Zhiyuan Zhang, Huajun Liu 等CVPR 2025
- BVT-IMA: Binary Vision Transformer with Information-Modified AttentionZhenyu Wang, Hao Luo, Xuemei Xie, Fan Wang 等AAAI 2024 · 被引用 4 次
