Precon: A Precision-Convertible Architecture for Accelerating Quantized Deep Learning Models across Various Domains Including LLMs
Jongwoo Park, Hyeonseong Kim, Jiyun Han, Seungkyu Choi
Abstract
The sensitivity of LLMs to quantization has driven the development of hardware accelerators tailored for specific low-precision configurations such as weight-only quantization and mixed-precision, which can introduce inefficiencies in dedicated hardware architecture. In this work, we propose Precon, a precision-convertible architecture designed to accelerate various quantized deep learning models, particularly LLMs, through a unified processing unit. By enabling on-the-fly switching between half-float (FP16) decoding and integer (INT) decomposition, the design effectively supports INT4-FP16, INT4-INT4, and INT4INT8 arithmetic within shared logic. Precon achieves up to speedup and 81.4% reduction in energy consumption compared to the baseline across various domains, including the support of both accurate and efficient acceleration of quantized LLMs.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get f0307208-3cbd-4fa2-8993-95e7c0853f82Related papers
- PacQ: A SIMT Microarchitecture for Efficient Dataflow in Hyper-asymmetric GEMMsRuokai Yin, Yuhang Li, Priyadarshini PandaDAC 2025 · 1 citation
- Linear Symmetric Quantization of Neural Networks for Low-precision Integer HardwareXiandong Zhao, Ying Wang, Xuyi Cai, Cheng Liu et al.ICLR 2020 · 67 citations
- Energy-Efficient and Dequantization-Free Quantization of LLMs: A Spiking Neural Network Approach to Salient Value MitigationChenyu Wang, Zhanglu Yan, Zhi Zhou, Xu Chen et al.WWW 2026
- Progressive Mixed-Precision Decoding for Efficient LLM InferenceHao Mark Chen, Fuwen Tan, Alexandros Kouris, Royson Lee et al.ICLR 2025 · 1 citation
- HAWQ-V3: Dyadic Neural Network QuantizationZhewei Yao, Zhen Dong, Zhangcheng Zheng, Amir Gholami et al.ICML 2021 · 240 citations
