BiT: Robustly Binarized Multi-distilled Transformer
Zechun Liu, Barlas Oguz, Aasish Pappu, Lin Xiao, Scott Yih, Meng Li, Raghuraman Krishnamoorthi, Yashar Mehdad
摘要
Modern pre-trained transformers have rapidly advanced the state-of-the-art in machine learning, but have also grown in parameters and computational complexity, making them increasingly difficult to deploy in resource-constrained environments. Binarization of the weights and activations of the network can significantly alleviate these issues, however, is technically challenging from an optimization perspective. In this work, we identify a series of improvements that enables binary transformers at a much higher accuracy than what was possible previously. These include a two-set binarization scheme, a novel elastic binary activation function with learned parameters, and a method to quantize a network to its limit by successively distilling higher precision models into lower precision students. These approaches allow for the first time, fully binarized transformer models that are at a practical level of accuracy, approaching a full-precision BERT baseline on the GLUE language understanding benchmark within as little as 5.9%. Code and models are available at: https://github.com/facebookresearch/bit . * Equal contribution Preprint. Under review.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper24
- PB-LLM: Partially Binarized Large Language ModelsZhihang Yuan, Yuzhang Shang, Zhen DongICLR 2024 · 被引用 91 次
- BiBench: Benchmarking and Analyzing Network BinarizationHaotong Qin, Mingyuan Zhang, Yifu Ding, Aoyu Li 等ICML 2023 · 被引用 53 次
- ParetoQ: Improving Scaling Laws in Extremely Low-bit LLM QuantizationZechun Liu, Changsheng Zhao, Hanxian Huang, Sijia Chen 等NeurIPS 2025 · 被引用 50 次
- SpikeLM: Towards General Spike-Driven Language Modeling via Elastic Bi-Spiking MechanismsXingrun Xing, Zheng Zhang, Ziyi Ni, Shitao Xiao 等ICML 2024 · 被引用 34 次
- ShiftAddViT: Mixture of Multiplication Primitives Towards Efficient Vision TransformerHaoran You, Huihong Shi, Yipin Guo, Yingyan LinNeurIPS 2023 · 被引用 27 次
它引用的顶会 Paper15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learned Step Size quantizationSteven K. Esser, Jeffrey L. McKinstry, Deepika Bablani, Rathinakumar Appuswamy 等ICLR 2020 · 被引用 1,037 次
- Q-BERT: Hessian Based Ultra Low Precision Quantization of BERTSheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma 等AAAI 2020 · 被引用 656 次
- Additive Powers-of-Two Quantization: An Efficient Non-uniform Discretization for Neural NetworksYuhang Li, Xin Dong, Wei WangICLR 2020 · 被引用 315 次
- Training with Quantization Noise for Extreme Model CompressionPierre Stock, Angela Fan, Benjamin Graham, Edouard Grave 等ICLR 2021 · 被引用 262 次
相关 Paper
- Understanding and Overcoming the Challenges of Efficient Transformer QuantizationYelysei Bondarenko, Markus Nagel, Tijmen BlankevoortEMNLP 2021 · 被引用 74 次
- TernaryBERT: Distillation-aware Ultra-low Bit BERTWei Zhang, Lu Hou, Yichun Yin, Lifeng Shang 等EMNLP 2020 · 被引用 147 次
- BinaryBERT: Pushing the Limit of BERT QuantizationHaoli Bai, Wei Zhang, Lu Hou, Lifeng Shang 等ACL 2021
- BiPFT: Binary Pre-trained Foundation Transformer with Low-Rank Estimation of Binarization Residual PolynomialsXingrun Xing, Li Du, Xinyuan Wang, Xianlin Zeng 等AAAI 2024 · 被引用 5 次
- Binary and Ternary Natural Language GenerationZechun Liu, Barlas Oguz, Aasish Pappu, Yangyang Shi 等ACL 2023 · 被引用 4 次
