Cambricon-P: A Bitflow Architecture for Arbitrary Precision Computing
Yifan Hao, Yongwei Zhao, Chenxiao Liu, Zidong Du, Shuyao Cheng, Xiaqing Li, Xing Hu, Qi Guo, Zhiwei Xu, Tianshi Chen
Abstract
Arbitrary precision computing (APC), where the digits vary from tens to millions of bits, is fundamental for scientific applications, such as mathematics, physics, chemistry, and biology. APC on existing platforms (e.g., CPUs and GPUs) is achieved by decomposing the original data into small pieces to accommodate to the low-bitwidth (e.g., 32-/64-bit) functional units. However, such fine-grained decomposition inevitably introduces large amounts of intermediates, bringing in intensive on-chip data traffic and long, complex dependency chains, so that causing low hardware utilization.To address this issue, we propose Cambricon-P, a bitflow architecture supporting monolithic large and flexible bitwidth operations for efficient APC processing, which avoids generating large amounts of intermediates from decomposition. Cambricon- P features a tightly-integrated computational architecture for processing different bitflows in parallel, where full bit-serial data paths are deployed. The bit-serial scheme still needs to eliminate the dependency chain of APC for exploiting parallelism within one monolithic large-bitwidth operation. For this purpose, Cambricon-P adopts a carry parallel computing mechanism, which enables recursively transforming the multiplication into smaller inner-products that can be performed in parallel between bit-indexed IPUs (Inner-Product Units). Furthermore, to improve the computing efficiency of APC, Cambricon- P employs a bit-indexed inner-product processing scheme, namely BIPS, to eliminate intra-IPU bit-level redundancy. Compared to Intel Xeon 6134 CPU, Cambricon-P achieves 100.98 performance on monolithic long multiplication, and 23.41/30.16 speedup and energy benefit over four real-world APC applications on average. Compared to NVidia V100 GPU, Cambricon-P also delivers the same throughput, as well as 430/60.5 lesser area and power, respectively, on batch-processing multiplications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 021552f8-0e89-4f5e-9269-6dae27c1cc61Builds on2
- CraterLake: a hardware accelerator for efficient unbounded computation on encrypted dataNikola Samardzic, Axel Feldmann, Aleksandar Krastev, Nathan Manohar et al.ISCA 2022 · 205 citations
- BTS: an accelerator for bootstrappable fully homomorphic encryptionSangpyo Kim, Jongmin Kim, Michael Jaemin Kim, Wonkyung Jung et al.ISCA 2022 · 184 citations
Related papers
- Cambricon-C: Efficient 4-Bit Matrix Unit via PrimitivizationYi Chen, Yongwei Zhao, Yifan Hao, Yuanbo Wen et al.MICRO 2024 · 8 citations
- Cambricon-D: Full-Network Differential Acceleration for Diffusion ModelsWeihao Kong, Yifan Hao, Qi Guo, Yongwei Zhao et al.ISCA 2024 · 26 citations
- Cambricon-U: A Systolic Random Increment Memory Architecture for Unary ComputingHongrui Guo, Yongwei Zhao, Zhangmai Li, Yifan Hao et al.MICRO 2023 · 2 citations
- FloatAP: Supporting High-Performance Floating-Point Arithmetic in Associative ProcessorsKailin Yang, José F. MartínezMICRO 2024 · 5 citations
- Cambricon-M: A Fibonacci-Coded Charge-Domain SRAM-Based CIM Accelerator for DNN InferenceHongrui Guo, Mo Zou, Yifan Hao, Zidong Du et al.MICRO 2024 · 3 citations
