Bucket Getter: A Bucket-based Processing Engine for Low-bit Block Floating Point (BFP) DNNs
Yun-Chen Lo, Ren-Shuo Liu
2023年份
9被引次数
4顶会引用
摘要
Block floating point (BFP), an efficient numerical system for deep neural networks (DNNs), achieves a good trade-off between dynamic range and hardware costs. Specifically, prior works have demonstrated that BFP format with 3 ∼ 5-bit mantissa can achieve FP32-comparable accuracy for various DNN workloads. We find that the floating-point adder (FP-Acc), which contains modules for normalization, alignment, addition, and fixed-point-to-floating-point (FXP2FP) conversion, dominates the power and area overheads, hence hindering the hardware efficiency of state-of-the-art low-bit BFP processing engines (BFP-PE).
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper4
- LUT Tensor Core: A Software-Hardware Co-Design for LUT-Based Low-Bit LLM InferenceZhiwen Mo, Lei Wang, Jianyu Wei, Zhichen Zeng 等ISCA 2025 · 被引用 17 次
- MX+: Pushing the Limits of Microscaling Formats for Efficient Large Language Model ServingJungi Lee, Junyong Park, Soohyun Cha, Jaehoon Cho 等MICRO 2025 · 被引用 7 次
- Exploring the Performance Improvement of Tensor Processing Engines through Transformation in the Bit-weight Dimension of MACsQizhe Wu, Huawen Liang, Yuchen Gui, Zhichen Zeng 等HPCA 2025 · 被引用 2 次
- M2XFP: A Metadata-Augmented Microscaling Data Format for Efficient Low-bit QuantizationWeiming Hu, Zihan Zhang, Haoyan Zhang, Chen Zhang 等ASPLOS 2026 · 被引用 2 次
相关 Paper
- DBPS: Dynamic Block Size and Precision Scaling for Efficient DNN Training Supported by RISC-V ISA ExtensionsSeunghyun Lee, Jeik Choi, Seock-Hwan Noh, Jahyun Koo 等DAC 2023 · 被引用 11 次
- FAST: DNN Training Under Variable Precision Block Floating Point with Stochastic RoundingSai Qian Zhang, Bradley McDanel, H. T. KungHPCA 2022 · 被引用 68 次
- A Block Minifloat Representation for Training Deep Neural NetworksSean Fox, Seyedramin Rasoulinezhad, Julian Faraone, David Boland 等ICLR 2021 · 被引用 31 次
- Winning Both the Accuracy of Floating Point Activation and the Simplicity of Integer ArithmeticYulhwa Kim, Jaeyong Jang, Jehun Lee, Jihoon Park 等ICLR 2023
- Distilling Bit-level Sparsity Parallelism for General Purpose Deep Learning AccelerationHang Lu, Liang Chang, Chenglong Li, Zixuan Zhu 等MICRO 2021 · 被引用 54 次
