Cambricon-C: Efficient 4-Bit Matrix Unit via Primitivization
Yi Chen, Yongwei Zhao, Yifan Hao, Yuanbo Wen, Yuntao Dai, Xiaqing Li, Yang Liu, Rui Zhang, Mo Zou, Xinkai Song, Xing Hu, Zidong Du
摘要
Deep learning trends to use low precision numeral formats to cope with the ever-growing model sizes. For example, the large language model LLaMA2 has been widely deployed in 4-bit precision. With larger models and fewer unique values caused by low precision, an increasing proportion of arithmetic in matrix multiplication is repeating. Although discussed in prior works, such value redundancy has not been fully exploited, and the cost to leverage the value redundancy often offsets any advantages. In this paper, we propose to primitivize the matrix multiplication, that is decomposing it down to the 1-ary successor function (a.k.a. counting) to merge repeating arithmetic. We revisited various techniques to propose Cambricon-C SA, a 4-bit primitive matrix multiplication unit that doubles the energy efficiency over conventional systolic arrays. Experimental results show that Cambricon-C SA can achieveenergy efficiency improvement compared with MAC-based systolic array.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- Carat: Unlocking Value-Level Parallelism for Multiplier-Free GEMMsZhewen Pan, Joshua San Miguel, Di WuASPLOS 2024 · 被引用 2 次
- Cambricon-P: A Bitflow Architecture for Arbitrary Precision ComputingYifan Hao, Yongwei Zhao, Chenxiao Liu, Zidong Du 等MICRO 2022 · 被引用 9 次
- Cambricon-U: A Systolic Random Increment Memory Architecture for Unary ComputingHongrui Guo, Yongwei Zhao, Zhangmai Li, Yifan Hao 等MICRO 2023 · 被引用 2 次
- uSystolic: Byte-Crawling Unary Systolic ArrayDi Wu, Joshua San MiguelHPCA 2022 · 被引用 28 次
- UGEMM: Unary Computing Architecture for GEMM ApplicationsDi Wu, Jingjie Li, Ruokai Yin, Hsuan Hsiao 等ISCA 2020 · 被引用 67 次
