Carat: Unlocking Value-Level Parallelism for Multiplier-Free GEMMs
Zhewen Pan, Joshua San Miguel, Di Wu
Abstract
In recent years, hardware architectures optimized for general matrix multiplication (GEMM) have been well studied to deliver better performance and efficiency for deep neural networks. With trends towards batched, low-precision data, e.g., FP8 format in this work, we observe that there is growing untapped potential for value reuse. We propose a novel computing paradigm, value-level parallelism, whereby unique products are computed only once, and different inputs subscribe to (select) their products via temporal coding. Our architecture, Carat, employs value-level parallelism and transforms multiplication into accumulation, performing GEMMs with efficient multiplier-free hardware. Experiments show that, on average, Carat improves iso-area throughput and energy efficiency by 1.02× and 1.06× over a systolic array and 3.2× and 4.3× when scaled up to multiple nodes.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get cab16bb5-8df8-43c9-a9a0-2969f19b236dCited by top-tier papers1
Ask how each one uses itRelated papers
- uSystolic: Byte-Crawling Unary Systolic ArrayDi Wu, Joshua San MiguelHPCA 2022 · 28 citations
- UGEMM: Unary Computing Architecture for GEMM ApplicationsDi Wu, Jingjie Li, Ruokai Yin, Hsuan Hsiao et al.ISCA 2020 · 67 citations
- Cambricon-C: Efficient 4-Bit Matrix Unit via PrimitivizationYi Chen, Yongwei Zhao, Yifan Hao, Yuanbo Wen et al.MICRO 2024 · 8 citations
- LUTein: Dense-Sparse Bit-Slice Architecture With Radix-4 LUT-Based Slice-Tensor Processing UnitsDongseok Im, Hoi-Jun YooHPCA 2024 · 10 citations
- SIMD2: a generalized matrix instruction set for accelerating tensor computation beyond GEMMYunan Zhang, Po-An Tsai, Hung-Wei TsengISCA 2022 · 6 citations
