Lune

DAC2025顶会

An Efficient Bit-level Sparse MAC-accelerated Architecture with SW/HW Co-design on FPGA

Chenming Zhang, Lei Gong, Chao Wang, Xuehai Zhou

2025年份
1被引次数

摘要

Exploring bit-level sparsity in the MAC process has been proven to be an important method for improving the efficiency of neural network feedforward processing. The reconfigurable platform offers possibilities for identifying the bitlevel unstructured redundancy during inference with different DNN models. Researchers noticed significant progress in valueaware accelerators on ASICs, yet we are concerned about the few studies on FPGAs. This paper observed the limitations of implementing bit-level sparsity optimizations using FPGA and proposed a software/architecture co-design solution. Specifically, by introducing LUT-friendly encoding with adaptable granularity and hardware structure supporting multiplication time uncertainty, we achieved a better trade-off between potential redundancy and accuracy with compatibility and scalability. Experiments show that under accurate calculation, PEs are up to 2.2×2.2 \times smaller than bit-parallel ones, and our design boosts performance by 1.04×1.04 \times to 1.74×1.74 \times and 1.40×1.40 \times to 2.79×2.79 \times over bitparallel and Booth-based designs, respectively.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖