uSystolic: Byte-Crawling Unary Systolic Array
Di Wu, Joshua San Miguel
Abstract
General matrix multiply (GEMM) is an important operation in broad applications, especially the thriving deep neural networks. To achieve low power consumption for GEMM, researchers have already leveraged unary computing, which manipulates bitstreams with extremely simple logic. However, existing unary architectures are not well generalizable to varying GEMM configurations in versatile applications and incompatible to the binary computing stack, imposing challenges to execute unary GEMM effortlessly. In this work, we address the problem by architecting a hybrid unary-binary systolic array, uSystolic, to inherit the legacy-binary data scheduling with slow (thus power-efficient) data movement, i.e., data bytes are crawling out from memory to drive uSystolic. uSystolic exhibits tremendous area and power improvements as a joint effect of 1) low-power computing kernel, 2) spatial-temporal bitstream reuse, and 3) on-chip SRAM elimination. For the evaluated edge computing scenario, compared with the binary parallel design, the rated-coded uSystolic reduces the systolic array area and total on-chip area by 59.0% and 91.3%, with the on-chip energy and power efficiency improved by up to 112.2× and 44.8× for AlexNet.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers3
- uBrain: a unary brain computer interfaceDi Wu, Jingjie Li, Zhewen Pan, Younghyun Kim et al.ISCA 2022 · 20 citations
- Cambricon-U: A Systolic Random Increment Memory Architecture for Unary ComputingHongrui Guo, Yongwei Zhao, Zhangmai Li, Yifan Hao et al.MICRO 2023 · 2 citations
- Mugi: Value Level Parallelism For Efficient LLMsDaniel Price, Prabhu Vellaisamy, John Paul Shen, Di WuASPLOS 2026
Related papers
- UGEMM: Unary Computing Architecture for GEMM ApplicationsDi Wu, Jingjie Li, Ruokai Yin, Hsuan Hsiao et al.ISCA 2020 · 67 citations
- Carat: Unlocking Value-Level Parallelism for Multiplier-Free GEMMsZhewen Pan, Joshua San Miguel, Di WuASPLOS 2024 · 2 citations
- Mix-GEMM: An efficient HW-SW Architecture for Mixed-Precision Quantized Deep Neural Networks Inference on Edge DevicesEnrico Reggiani, Alessandro Pappalardo, Max Doblas, Miquel Moretó et al.HPCA 2023 · 31 citations
- Cambricon-C: Efficient 4-Bit Matrix Unit via PrimitivizationYi Chen, Yongwei Zhao, Yifan Hao, Yuanbo Wen et al.MICRO 2024 · 8 citations
- NCPU: An Embedded Neural CPU Architecture on Resource-Constrained Low Power Devices for Real-time End-to-End PerformanceTianyu Jia, Yuhao Ju, Russ Joseph, Jie GuMICRO 2020 · 21 citations
