Cambricon-U: A Systolic Random Increment Memory Architecture for Unary Computing
Hongrui Guo, Yongwei Zhao, Zhangmai Li, Yifan Hao, Chang Liu, Xinkai Song, Xiaqing Li, Zidong Du, Rui Zhang, Qi Guo, Tianshi Chen, Zhiwei Xu
摘要
Unary computing, whose arithmetics require only one logic gate, has enabled efficient DNN processing, especially on strictly powerconstrained devices. However, unary computing still confronts the power efficiency bottleneck for buffering unary bitstreams. The buffering of unary bitstreams requires accumulating bits into large bitwidth binary numbers. The large bitwidth binary number needs to activate all bits per cycle in case of carry propagation. As a result, the accumulation process accounts for 32%-70% of the power budget.
To push the boundary of power efficiency, we propose Cambricon-U, a systolic random increment memory architecture featuring efficient accumulation. By leveraging skew number data format, Cambricon-U only activates no more than three bits (instead of all bits) from each number per accumulating cycle. Experimental results show that Cambricon-U reduces 51% power on unary accumulation, and improves 1.18-1.45× energy efficiency over uSystolic, the SOTA unary computing scheme baseline, with -1.9%∼+0.77% area overhead.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Hardwired-Neuron Language Processing Units as General-Purpose Cognitive SubstratesYang Liu, Yi Chen, Yongwei Zhao, Yifan Hao 等ASPLOS 2026
- Mugi: Value Level Parallelism For Efficient LLMsDaniel Price, Prabhu Vellaisamy, John Paul Shen, Di WuASPLOS 2026
它引用的顶会 Paper4
- UGEMM: Unary Computing Architecture for GEMM ApplicationsDi Wu, Jingjie Li, Ruokai Yin, Hsuan Hsiao 等ISCA 2020 · 被引用 67 次
- uSystolic: Byte-Crawling Unary Systolic ArrayDi Wu, Joshua San MiguelHPCA 2022 · 被引用 28 次
- Cambricon-Q: A Hybrid Architecture for Efficient TrainingYongwei Zhao, Chang Liu, Zidong Du, Qi Guo 等ISCA 2021 · 被引用 28 次
- uBrain: a unary brain computer interfaceDi Wu, Jingjie Li, Zhewen Pan, Younghyun Kim 等ISCA 2022 · 被引用 20 次
相关 Paper
- Cambricon-M: A Fibonacci-Coded Charge-Domain SRAM-Based CIM Accelerator for DNN InferenceHongrui Guo, Mo Zou, Yifan Hao, Zidong Du 等MICRO 2024 · 被引用 3 次
- Cambricon-C: Efficient 4-Bit Matrix Unit via PrimitivizationYi Chen, Yongwei Zhao, Yifan Hao, Yuanbo Wen 等MICRO 2024 · 被引用 8 次
- Comparison-Free Bit-Stream Generation for Cost-Efficient Unary ComputingFaeze S. Banitaba, Amir Hossein Jalilvand, M. Hassan Najafi, Sercan AygunDAC 2025
- Cambricon-P: A Bitflow Architecture for Arbitrary Precision ComputingYifan Hao, Yongwei Zhao, Chenxiao Liu, Zidong Du 等MICRO 2022 · 被引用 9 次
- Cambricon-D: Full-Network Differential Acceleration for Diffusion ModelsWeihao Kong, Yifan Hao, Qi Guo, Yongwei Zhao 等ISCA 2024 · 被引用 26 次
