A Block Minifloat Representation for Training Deep Neural Networks
Sean Fox, Seyedramin Rasoulinezhad, Julian Faraone, David Boland, Philip H. W. Leong
Abstract
Training Deep Neural Networks (DNN) with high efficiency can be difficult to achieve with native floating-point representations and commercially available hardware. Specialized arithmetic with custom acceleration offers perhaps the most promising alternative. Ongoing research is trending towards narrow floating-point representations, called minifloats, that pack more operations for a given silicon area and consume less power. In this paper, we introduce Block Minifloat (BM), a new spectrum of minifloat formats capable of training DNNs end-to-end with only 4-8 bit weight, activation and gradient tensors. While standard floating-point representations have two degrees of freedom, via the exponent and mantissa, BM exposes the exponent bias as an additional field for optimization. Crucially, this enables training with fewer exponent bits, yielding dense integer-like hardware for fused multiply-add (FMA) operations. For ResNet trained on ImageNet, 6-bit BM achieves almost no degradation in floating-point accuracy with FMA units that are smaller and consume less energy than FP8 (FP32). Furthermore, our 8-bit BM format matches floating-point accuracy while delivering a higher computational density and faster expected training times.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 03de311c-7834-4fa1-a3c7-785041720f2dCited by top-tier papers4
- LQER: Low-Rank Quantization Error Reconstruction for LLMsCheng Zhang, Jianyi Cheng, George Anthony Constantinides, Yiren ZhaoICML 2024 · 33 citations
- Revisiting Block-based Quantisation: What is Important for Sub-8-bit LLM Inference?Cheng Zhang, Jianyi Cheng, Ilia Shumailov, George A. Constantinides et al.EMNLP 2023 · 5 citations
- BBAL: A Bidirectional Block Floating Point-Based Quantisation Accelerator for Large Language ModelsXiaomeng Han, Yuan Cheng, Jing Wang, Junyang Lu et al.DAC 2025 · 5 citations
- FF-INT8: Efficient Forward-Forward DNN Training on Edge Devices with INT8 PrecisionJingxiao Ma, Priyadarshini Panda, Sherief RedaDAC 2025 · 1 citation
Related papers
- FAST: DNN Training Under Variable Precision Block Floating Point with Stochastic RoundingSai Qian Zhang, Bradley McDanel, H. T. KungHPCA 2022 · 68 citations
- DBPS: Dynamic Block Size and Precision Scaling for Efficient DNN Training Supported by RISC-V ISA ExtensionsSeunghyun Lee, Jeik Choi, Seock-Hwan Noh, Jahyun Koo et al.DAC 2023 · 11 citations
- Bucket Getter: A Bucket-based Processing Engine for Low-bit Block Floating Point (BFP) DNNsYun-Chen Lo, Ren-Shuo LiuMICRO 2023 · 9 citations
- Shifted and Squeezed 8-bit Floating Point format for Low-Precision Training of Deep Neural NetworksLéopold Cambier, Anahita Bhiwandiwalla, Ting Gong, Oguz H. Elibol et al.ICLR 2020 · 53 citations
- Distilling Bit-level Sparsity Parallelism for General Purpose Deep Learning AccelerationHang Lu, Liang Chang, Chenglong Li, Zixuan Zhu et al.MICRO 2021 · 54 citations
