BitMoD: Bit-serial Mixture-of-Datatype LLM Acceleration
Yuzong Chen, Ahmed F. AbouElhamayed, Xilai Dai, Yang Wang, Marta Andronic, George A. Constantinides, Mohamed S. Abdelfattah
Abstract
Large language models (LLMs) have demonstrated remarkable performance across various machine learning tasks. Yet the substantial memory footprint of LLMs significantly hinders their deployment. In this paper, we improve the accessibility of LLMs through BitMoD 1 , an algorithm-hardware co-design solution that enables efficient LLM acceleration at low weight precision. On the algorithm side, BitMoD introduces fine-grained data type adaptation that uses a different numerical data type to quantize a group of (e.g., 128) weights. Through the careful design of these new data types, BitMoD is able to quantize LLM weights to very low precision (e.g., 4 bits and 3 bits) while maintaining high accuracy. On the hardware side, BitMoD employs a bitserial processing element to easily support multiple numerical precisions and data types; our hardware design includes two key innovations: First, it employs a unified representation to process different weight data types, thus reducing the hardware cost. Second, it adopts a bit-serial dequantization unit to rescale the per-group partial sum with minimal hardware overhead. Our evaluation on six representative LLMs demonstrates that BitMoD significantly outperforms state-of-the-art LLM quantization and acceleration methods. For discriminative tasks, BitMoD can quantize LLM weights to 4-bit with < 0.5% accuracy loss on average. For generative tasks, BitMoD is able to quantize LLM weights to 3-bit while achieving better perplexity than prior LLM quantization scheme. Combining the superior model performance with an efficient accelerator design, BitMoD achieves an average of 1.69× and 1.48× speedups compared to prior LLM accelerators ANT and OliVe, respectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d06d8d3e-4001-4d39-a357-a1da9dd83ba4Cited by top-tier papers7
- -LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical FormatsYuzong Chen, Chao Fang, Xilai Dai, Yuheng Wu et al.ISCA 2026 · 4 citations
- PADE: A Predictor-Free Sparse Attention Accelerator via Unified Execution and Stage FusionHuizheng Wang, Hongbin Wang, Zichuan Wang, Zhiheng Yue et al.HPCA 2026 · 2 citations
- GyRot: Leveraging Hidden Synergy Between Rotation and Fine-Grained Group Quantization for Low-Bit LLM InferenceSangjin Kim, Yuseon Chou, Byeongcheol Kim, Jungjun Oh et al.HPCA 2026 · 2 citations
- M2XFP: A Metadata-Augmented Microscaling Data Format for Efficient Low-bit QuantizationWeiming Hu, Zihan Zhang, Haoyan Zhang, Chen Zhang et al.ASPLOS 2026 · 2 citations
- Focus: A Streaming Concentration Architecture for Efficient Vision-Language ModelsChiyue Wei, Cong Guo, Junyao Zhang, Haoxuan Shan et al.HPCA 2026 · 2 citations
Builds on26
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu et al.ICML 2023 · 1,493 citations
Related papers
- An Algorithm-Hardware Co-design Based on Revised Microscaling Format Quantization for Accelerating Large Language ModelsYingbo Hao, Huangxu Chen, Yi Zou, Yanfeng YangDAC 2025 · 1 citation
- Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data FormatChao Fang, Man Shi, Robin Geens, Arne Symons et al.HPCA 2025 · 15 citations
- Amove: Accelerating LLMs through Mitigating Outliers and Salient Points via Fine-Grained Grouped Vectorized Data TypeXilong Xie, Liang Wang, Limin Xiao, Meng Han et al.MICRO 2025 · 2 citations
- S-Quant: Rethinking Weight Quantization with Seed-Based GenerationMingzi Wang, Lancheng Zou, Shuo Yin, Zhuolun He et al.ICML 2026
- OPAL: Outlier-Preserved Microscaling Quantization Accelerator for Generative Large Language ModelsJahyun Koo, Dahoon Park, Sangwoo Jung, Jaeha KungDAC 2024 · 12 citations
