April: Accuracy-Improved Floating-Point Approximation For Neural Network Accelerators
Yonghao Chen, Jiaxiang Zou, Xinyu Chen
摘要
Neural Networks (NNs) have achieved breakthroughs in computer vision and natural language processing. However, modern models are computationally expensive, with floating-point operations posing a major bottleneck. Floatingpoint approximation, such as Mitchell’s logarithm, enables floating-point multiplication using simpler integer additions, thereby improving hardware efficiency. However, its practical adoption is hindered by challenges such as precision degradation, efficient hardware integrations, and management of trade-offs between accuracy and resource efficiency. In this paper, we propose a hardware-efficient down-samplingbased compensation method to mitigate precision loss and a flexible bias mechanism to accommodate diverse data distributions in NN models. Building on this foundation, we design configurable systolic arrays optimized for NN accelerators. To further support practical adoption, we introduce April, a co-design framework that balances the accuracy and resource usage of generated synthesizable systolic arrays. Our FPGA-based evaluations demonstrate that April-generated systolic arrays reduce root mean square error (RMSE) by up to and achieve area reduction even compared to INT8-based implementations while maintaining comparable or improved model accuracy. Our design is open-sourced at https://github.com/CLabGit/April
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- UniCore: A Bit-Width Scalable GEMM Unit for Unified LLM InferenceYonghao Chen, Jiaxiang Zou, Xingyu Chen, Chenxi Xu 等ISCA 2026
- MixFP4: Enhancing NVFP4 with Adaptive FP4/INT4 Block RepresentationsJiaxiang Zou, Yonghao Chen, Ruilong WU, Xinyu ChenICML 2026
相关 Paper
- M2XFP: A Metadata-Augmented Microscaling Data Format for Efficient Low-bit QuantizationWeiming Hu, Zihan Zhang, Haoyan Zhang, Chen Zhang 等ASPLOS 2026 · 被引用 2 次
- FlexiQ: Adaptive Mixed-Precision Quantization for Latency/Accuracy Trade-Offs in Deep Neural NetworksJaemin Kim, Hongjun Um, Sungkyun Kim, Yongjun Park 等EuroSys 2026
- A2Q: Accumulator-Aware Quantization with Guaranteed Overflow AvoidanceIan Colbert, Alessandro Pappalardo, Jakoba Petri-KoenigICCV 2023 · 被引用 19 次
- DUET: Boosting Deep Neural Network Efficiency on Dual-Module ArchitectureLiu Liu, Zheng Qu, Lei Deng, Fengbin Tu 等MICRO 2020 · 被引用 27 次
- CVMAX: Accelerator Architecture with Polar Form Multiplication for Complex-Valued Neural NetworksHyunwuk Lee, Sungbin Kim, Sungwoo Kim, Won Woo RoDAC 2025 · 被引用 1 次
