OliVe: Accelerating Large Language Models via Hardware-friendly Outlier-Victim Pair Quantization
Cong Guo, Jiaming Tang, Weiming Hu, Jingwen Leng, Chen Zhang, Fan Yang, Yunxin Liu, Minyi Guo, Yuhao Zhu
Abstract
Transformer-based large language models (LLMs) have achieved great success with the growing model size. LLMs' size grows by 240× every two years, which outpaces the hardware progress and makes model inference increasingly costly. Model quantization is a promising approach to mitigate the widening gap between LLM size and hardware capacity. However, the existence of outliers, values with significant magnitudes, in LLMs makes existing quantization methods less effective. Prior outlier-aware quantization schemes adopt sparsity encoding techniques to separate outliers from normal values where the process requires global coordination (e.g., a global sparsity coordination list). This incurs complex encoding/decoding hardware logics and an extra orchestration controller for the computation between outlier and normal values. As such, it is not hardware-efficient and hence only achieves sub-optimal quantization benefits.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ccf3ada2-2c16-47cf-91ee-05537e7c24ceCited by top-tier papers50
- PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model InferenceYufeng Gu, Alireza Khadem, Sumanth Umesh, Ning Liang et al.ASPLOS 2025 · 44 citations
- Tender: Accelerating Large Language Models via Tensor Decomposition and Runtime RequantizationJungi Lee, Wonbeom Lee, Jaewoong SimISCA 2024 · 41 citations
- Outlier Suppression+: Accurate quantization of large language models by equivalent and effective shifting and scalingXiuying Wei, Yunchen Zhang, Yuhang Li, Xiangguo Zhang et al.EMNLP 2023 · 40 citations
- Duplex: A Device for Large Language Models with Mixture of Experts, Grouped Query Attention, and Continuous BatchingSungmin Yun, Kwanhee Kyung, Juhwan Cho, Jaewan Choi et al.MICRO 2024 · 40 citations
- PAPI: Exploiting Dynamic Parallelism in Large Language Model Decoding with a Processing-In-Memory-Enabled Computing SystemYintao He, Haiyu Mao, Christina Giannoula, Mohammad Sadrosadati et al.ASPLOS 2025 · 37 citations
Builds on26
- Q-BERT: Hessian Based Ultra Low Precision Quantization of BERTSheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma et al.AAAI 2020 · 656 citations
- HAWQ: Hessian AWare Quantization of Neural Networks With Mixed-PrecisionZhen Dong, Zhewei Yao, Amir Gholami, Michael W. Mahoney et al.ICCV 2019 · 645 citations
- Ansor: Generating High-Performance Tensor Programs for Deep LearningLianmin Zheng, Chengfan Jia, Minmin Sun, Zhao Wu et al.OSDI 2020 · 551 citations
- SIGMA: A Sparse and Irregular GEMM Accelerator with Flexible Interconnects for DNN TrainingEric Qin, Ananda Samajdar, Hyoukjun Kwon, Vineet Nadella et al.HPCA 2020 · 490 citations
- HAWQ-V2: Hessian Aware trace-Weighted Quantization of Neural NetworksZhen Dong, Zhewei Yao, Daiyaan Arfeen, Amir Gholami et al.NeurIPS 2020 · 434 citations
Related papers
- DuoQ: A DSP Utilization-aware and Outlier-free Quantization for FPGA-based LLMs AccelerationZhuoquan Yu, Huidong Ji, Yue Cao, Junfu Wu et al.DAC 2025 · 1 citation
- Compressing Large Language Models by Joint Sparsification and QuantizationJinyang Guo, Jianyu Wu, Zining Wang, Jiaheng Liu et al.ICML 2024 · 33 citations
- Oltron: Algorithm-Hardware Co-design for Outlier-Aware Quantization of LLMs with Inter-/Intra-Layer AdaptationChenhao Xue, Chen Zhang, Xun Jiang, Zhutianya Gao et al.DAC 2024 · 11 citations
- DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMsHaokun Lin, Haobo Xu, Yichen Wu, Jingzhi Cui et al.NeurIPS 2024 · 206 citations
- OutlierCIM: Outlier-Aware Digital CIM-Based LLM Accelerator with Hybrid-Strategy Quantization and Unified FP-INT ComputationZihan Zou, Shikuang Chen, Chen Zhang, Xing Wang et al.DAC 2025
