Mix-GEMM: An efficient HW-SW Architecture for Mixed-Precision Quantized Deep Neural Networks Inference on Edge Devices
Enrico Reggiani, Alessandro Pappalardo, Max Doblas, Miquel Moretó, Mauro Olivieri, Osman Sabri Unsal, Adrián Cristal
摘要
Deep Neural Network (DNN) inference based on quantized narrow-precision integer data represents a promising research direction toward efficient deep learning computations on edge and mobile devices. On one side, recent progress of Quantization-Aware Training (QAT) frameworks aimed at improving the accuracy of extremely quantized DNNs allows achieving results close to Floating-Point 32 (FP32), and provides high flexibility concerning the data sizes selection. Unfortunately, current Central Processing Unit (CPU) architectures and Instruction Set Architectures (ISAs) targeting resource-constrained devices present limitations on the range of data sizes supported to compute DNN kernels.This paper presents Mix-GEMM, a hardware-software co-designed architecture capable of efficiently computing quantized DNN convolutional kernels based on byte and sub-byte data sizes. Mix-GEMM accelerates General Matrix Multiplication (GEMM), representing the core kernel of DNNs, supporting all data size combinations from 8- to 2-bit, including mixed-precision computations, and featuring performance that scale with the decreasing of the computational data sizes. Our experimental evaluation, performed on representative quantized Convolutional Neural Networks (CNNs), shows that a RISC-V based edge System-on-Chip (SoC) integrating Mix-GEMM achieves up to 1.3 TOPS/W in energy efficiency, and up to 13.6 GOPS in throughput, gaining from 5.3× to 15.1× in performance over the OpenBLAS GEMM frameworks running on a commercial RISC-V based edge processor. By performing synthesis and Place and Route (PnR) of the enhanced SoC in Global Foundries 22nm FDX technology, we show that Mix-GEMM only accounts for 1% of the overall area consumption.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper5
- M-ANT: Efficient Low-bit Group Quantization for LLMs via Mathematically Adaptive Numerical TypeWeiming Hu, Haoyan Zhang, Cong Guo, Yu Feng 等HPCA 2025 · 被引用 19 次
- Oaken: Fast and Efficient LLM Serving with Online-Offline Hybrid KV Cache QuantizationMinsu Kim, Seongmin Hong, Ryeowook Ko, Soongyu Choi 等ISCA 2025 · 被引用 17 次
- JAQ: Joint Efficient Architecture Design and Low-Bit Quantization with Hardware-Software Co-ExplorationMingzi Wang, Yuan Meng, Chen Tang, Weixiang Zhang 等AAAI 2025 · 被引用 3 次
- Empowering Vector Architectures for ML: The CAMP Architecture for Matrix MultiplicationMohammadreza Esmali Nojehdeh, Hossein Mokhtarnia, Julian Pavon, Narcís Rodas 等MICRO 2025 · 被引用 1 次
- PacQ: A SIMT Microarchitecture for Efficient Dataflow in Hyper-asymmetric GEMMsRuokai Yin, Yuhang Li, Priyadarshini PandaDAC 2025 · 被引用 1 次
相关 Paper
- BiSon-e: a lightweight and high-performance accelerator for narrow integer linear algebra computing on the edgeEnrico Reggiani, Cristóbal Ramírez Lazo, Roger Figueras Bagué, Adrián Cristal 等ASPLOS 2022 · 被引用 11 次
- NCPU: An Embedded Neural CPU Architecture on Resource-Constrained Low Power Devices for Real-time End-to-End PerformanceTianyu Jia, Yuhao Ju, Russ Joseph, Jie GuMICRO 2020 · 被引用 21 次
- BiQGEMM: matrix multiplication with lookup table for binary-coding-based quantized DNNsYongkweon Jeon, Baeseong Park, Se Jung Kwon, Byeongwook Kim 等SC 2020 · 被引用 31 次
- Mix and Match: A Novel FPGA-Centric Deep Neural Network Quantization FrameworkSung-En Chang, Yanyu Li, Mengshu Sun, Runbin Shi 等HPCA 2021 · 被引用 125 次
- HiPACK: Efficient Sub-8-Bit Direct Convolution with SIMD and Bitwise ManagementYao Chen, Cheng Gong, Bingsheng HeMICRO 2025 · 被引用 2 次
