MERSIT: A Hardware-Efficient 8-bit Data Format with Enhanced Post-Training Quantization DNN Accuracy
Nguyen-Dong Ho, Gyujun Jeong, Cheol-Min Kang, Seungkyu Choi, Ik-Joon Chang
Abstract
Post-training quantization (PTQ) models utilizing conventional 8-bit Integer or floating-point formats still exhibit significant accuracy drops in modern deep neural networks (DNNs), rendering them unreliable. This paper presents MERSIT, a novel 8-bit PTQ data format designed for various DNNs. While leveraging the dynamic configuration of exponent and fraction bits derived from Posit data format, MERSIT demonstrates enhanced hardware efficiency through the proposed merged decoding scheme. Our evaluation indicates that MERSIT yields more reliable 8-bit PTQ models, exhibiting superior accuracy across various DNNs compared to conventional floating-point formats. Furthermore, the proposed processing unit saves 26.6% in area and 22.2% in power consumption compared to the Posit-based unit, while maintaining comparable efficiency to the floating-point-based unit.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get bfc9ebb8-fff0-45a4-9ac0-3fd3e884adf1Related papers
- FP8 Quantization: The Power of the ExponentAndrey Kuzmin, Mart van Baalen, Yuwei Ren, Markus Nagel et al.NeurIPS 2022 · 154 citations
- BRECQ: Pushing the Limit of Post-Training Quantization by Block ReconstructionYuhang Li, Ruihao Gong, Xu Tan, Yang Yang et al.ICLR 2021 · 619 citations
- DQT: Dynamic Quantization Training via Dequantization-Free Nested Integer ArithmeticHazem Hesham Yousef Shalby, Fabrizio Pittorino, Francesca Palermo, Diana Trojaniello et al.AAAI 2026 · 2 citations
- Shifted and Squeezed 8-bit Floating Point format for Low-Precision Training of Deep Neural NetworksLéopold Cambier, Anahita Bhiwandiwalla, Ting Gong, Oguz H. Elibol et al.ICLR 2020 · 53 citations
- Post-Training Sparsity-Aware QuantizationGil Shomron, Freddy Gabbay, Samer Kurzum, Uri C. WeiserNeurIPS 2021 · 47 citations
