uSystolic: Byte-Crawling Unary Systolic Array
Di Wu, Joshua San Miguel
摘要
General matrix multiply (GEMM) is an important operation in broad applications, especially the thriving deep neural networks. To achieve low power consumption for GEMM, researchers have already leveraged unary computing, which manipulates bitstreams with extremely simple logic. However, existing unary architectures are not well generalizable to varying GEMM configurations in versatile applications and incompatible to the binary computing stack, imposing challenges to execute unary GEMM effortlessly. In this work, we address the problem by architecting a hybrid unary-binary systolic array, uSystolic, to inherit the legacy-binary data scheduling with slow (thus power-efficient) data movement, i.e., data bytes are crawling out from memory to drive uSystolic. uSystolic exhibits tremendous area and power improvements as a joint effect of 1) low-power computing kernel, 2) spatial-temporal bitstream reuse, and 3) on-chip SRAM elimination. For the evaluated edge computing scenario, compared with the binary parallel design, the rated-coded uSystolic reduces the systolic array area and total on-chip area by 59.0% and 91.3%, with the on-chip energy and power efficiency improved by up to 112.2× and 44.8× for AlexNet.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper3
- uBrain: a unary brain computer interfaceDi Wu, Jingjie Li, Zhewen Pan, Younghyun Kim 等ISCA 2022 · 被引用 20 次
- Cambricon-U: A Systolic Random Increment Memory Architecture for Unary ComputingHongrui Guo, Yongwei Zhao, Zhangmai Li, Yifan Hao 等MICRO 2023 · 被引用 2 次
- Mugi: Value Level Parallelism For Efficient LLMsDaniel Price, Prabhu Vellaisamy, John Paul Shen, Di WuASPLOS 2026
相关 Paper
- UGEMM: Unary Computing Architecture for GEMM ApplicationsDi Wu, Jingjie Li, Ruokai Yin, Hsuan Hsiao 等ISCA 2020 · 被引用 67 次
- Carat: Unlocking Value-Level Parallelism for Multiplier-Free GEMMsZhewen Pan, Joshua San Miguel, Di WuASPLOS 2024 · 被引用 2 次
- Mix-GEMM: An efficient HW-SW Architecture for Mixed-Precision Quantized Deep Neural Networks Inference on Edge DevicesEnrico Reggiani, Alessandro Pappalardo, Max Doblas, Miquel Moretó 等HPCA 2023 · 被引用 31 次
- Cambricon-C: Efficient 4-Bit Matrix Unit via PrimitivizationYi Chen, Yongwei Zhao, Yifan Hao, Yuanbo Wen 等MICRO 2024 · 被引用 8 次
- NCPU: An Embedded Neural CPU Architecture on Resource-Constrained Low Power Devices for Real-time End-to-End PerformanceTianyu Jia, Yuhao Ju, Russ Joseph, Jie GuMICRO 2020 · 被引用 21 次
