Mokey: enabling narrow fixed-point inference for out-of-the-box floating-point transformer models
Ali Hadi Zadeh, Mostafa Mahmoud, Ameer Abdelhadi, Andreas Moshovos
摘要
Increasingly larger and better Transformer models keep advancing state-of-the-art accuracy and capability for Natural Language Processing applications. These models demand more computational power, storage, and energy. Mokey reduces the footprint of state-of-the-art 32-bit or 16-bit floating-point transformer models by quantizing all values to 4-bit indexes into dictionaries of representative 16-bit fixed-point centroids.
Mokey does not need fine-tuning, an essential feature as often the training resources or datasets are not available to many. Exploiting the range of values that naturally occur in transformer models, Mokey selects centroid values to also fit an exponential curve. This unique feature enables Mokey to replace the bulk of the original multiply-accumulate operations with narrow 3b fixed-point additions resulting in an area-and energy-efficient hardware accelerator design. Over a set of state-of-the-art transformer models, the Mokey accelerator delivers an order of magnitude improvements in energy efficiency over a Tensor Cores-based accelerator while improving performance by at least 4× and as much as 15× depending on the model and on-chip buffering capacity. Optionally, Mokey can be used as a memory compression assist for any other accelerator, transparently stashing wide floating-point or fixed-point activations or weights into narrow 4-bit indexes. Mokey proves superior to prior state-ofthe-art quantization methods for Transformers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Duplex: A Device for Large Language Models with Mixture of Experts, Grouped Query Attention, and Continuous BatchingSungmin Yun, Kwanhee Kyung, Juhwan Cho, Jaewan Choi 等MICRO 2024 · 被引用 40 次
- BitMoD: Bit-serial Mixture-of-Datatype LLM AccelerationYuzong Chen, Ahmed F. AbouElhamayed, Xilai Dai, Yang Wang 等HPCA 2025 · 被引用 23 次
- M-ANT: Efficient Low-bit Group Quantization for LLMs via Mathematically Adaptive Numerical TypeWeiming Hu, Haoyan Zhang, Cong Guo, Yu Feng 等HPCA 2025 · 被引用 19 次
- LUT Tensor Core: A Software-Hardware Co-Design for LUT-Based Low-Bit LLM InferenceZhiwen Mo, Lei Wang, Jianyu Wei, Zhichen Zeng 等ISCA 2025 · 被引用 17 次
- Oaken: Fast and Efficient LLM Serving with Online-Offline Hybrid KV Cache QuantizationMinsu Kim, Seongmin Hong, Ryeowook Ko, Soongyu Choi 等ISCA 2025 · 被引用 17 次
它引用的顶会 Paper14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 被引用 3,729 次
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 被引用 2,665 次
相关 Paper
- QUQ: Quadruplet Uniform Quantization for Efficient Vision Transformer InferenceXinkuang Geng, Siting Liu, Leibo Liu, Jie Han 等DAC 2024 · 被引用 5 次
- 8-bit Transformer Inference and Fine-tuning for Edge AcceleratorsJeffrey Yu, Kartik Prabhu, Yonatan Urman, Robert M. Radway 等ASPLOS 2024 · 被引用 26 次
- GOBO: Quantizing Attention-Based NLP Models for Low Latency and Energy Efficient InferenceAli Hadi Zadeh, Isak Edo, Omar Mohamed Awad, Andreas MoshovosMICRO 2020 · 被引用 11 次
- Training Transformers with 4-bit IntegersHaocheng Xi, Changhao Li, Jianfei Chen, Jun ZhuNeurIPS 2023 · 被引用 96 次
- Algorithm-Hardware Co-Design of Adaptive Floating-Point Encodings for Resilient Deep Learning InferenceThierry Tambe, En-Yu Yang, Zishen Wan, Yuntian Deng 等DAC 2020 · 被引用 67 次
