Empowering Vector Architectures for ML: The CAMP Architecture for Matrix Multiplication
Mohammadreza Esmali Nojehdeh, Hossein Mokhtarnia, Julian Pavon, Narcís Rodas, Roger Figueras Bagué, Enrico Reggiani, Miquel Moretó, Osman S. Unsal, Adrián Cristal, Eduard Ayguadé
Abstract
This study presents the Cartesian Accumulative Matrix Pipeline (CAMP) architecture, a novel approach designed to enhance matrix multiplication in Vector Architectures (VAs) and Single Instruction Multiple Data (SIMD) units. CAMP improves processing efficiency of Quantized Neural Networks (QNNs).
Matrix multiplication is a cornerstone of machine learning applications, and its quantized versions are increasingly popular for more efficient operations. Unfortunately, existing VAs and SIMDsupport units struggle to efficiently handle these quantized formats. In this work, we propose CAMP, a simple yet effective architecture that leverages a hybrid multiplier. The CAMP architecture significantly advances the performance of vector architectures in handling quantized data, enabling more efficient execution of matrix multiplication across various platforms, specifically targeting the ARMv8 Scalable Vector Extension (SVE) and edge RISC-V with SIMD-support unit architecture. In addition to increasing throughput, CAMP's architectural design also contributes to energy efficiency, making it an effective solution for low-power applications.
Evaluations on a range of Large Language Models (LLMs) and Convolutional Neural Networks (CNNs) demonstrate that matrix multiplication operations using the proposed micro-architecture achieve up to 17× and 23× performance improvements compared to their respective baselines, the ARM A64FX core and a RISC-V-based edge System-on-Chip (SoC). Furthermore, synthesis and place-and-route (PnR) of the CAMP micro-architecture using Synopsys tools-targeting ARM TSMC 7nm for A64FX and Global-Foundries 22nm for the RISC-V SoC-add only 1% and 4% area overhead, respectively, compared to the baseline designs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on10
- FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPUYing Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li et al.ICML 2023 · 683 citations
- SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head PruningHanrui Wang, Zhekai Zhang, Song HanHPCA 2021 · 412 citations
- Mix-GEMM: An efficient HW-SW Architecture for Mixed-Precision Quantized Deep Neural Networks Inference on Edge DevicesEnrico Reggiani, Alessandro Pappalardo, Max Doblas, Miquel Moretó et al.HPCA 2023 · 31 citations
- DeepBurning-SEG: Generating DNN Accelerators of Segment-Grained Pipeline ArchitectureXuyi Cai, Ying Wang, Xiaohan Ma, Yinhe Han et al.MICRO 2022 · 25 citations
- DiVa: An Accelerator for Differentially Private Machine LearningBeomsik Park, Ranggi Hwang, Dongho Yoon, Yoonhyuk Choi et al.MICRO 2022 · 12 citations
Related papers
- Exploring the Performance Improvement of Tensor Processing Engines through Transformation in the Bit-weight Dimension of MACsQizhe Wu, Huawen Liang, Yuchen Gui, Zhichen Zeng et al.HPCA 2025 · 2 citations
- BiQGEMM: matrix multiplication with lookup table for binary-coding-based quantized DNNsYongkweon Jeon, Baeseong Park, Se Jung Kwon, Byeongwook Kim et al.SC 2020 · 31 citations
- HiPACK: Efficient Sub-8-Bit Direct Convolution with SIMD and Bitwise ManagementYao Chen, Cheng Gong, Bingsheng HeMICRO 2025 · 2 citations
- XtraMAC: An Efficient MAC Architecture for Mixed-Precision LLM Inference on FPGAFeng Yu, Hongshi Tan, Yao Chen, Weng-Fai Wong et al.ISCA 2026
- uSystolic: Byte-Crawling Unary Systolic ArrayDi Wu, Joshua San MiguelHPCA 2022 · 28 citations
