Be Like Water: Adaptive Floating Point for Machine Learning
Thomas Y. Yeh, Max Sterner, Zerlina Lai, Brandon Chuang, Alexander Ihler
摘要
In the pursuit of optimizing memory and compute density to accelerate machine learning applications, reduced precision training and inference has been an active area of research. While some approaches selectively apply low precision computations, this may require costly off-chip data transfers or mixed precision support. In this paper, we propose a novel numerical representation, Adaptive Floating Point (AFP), that dynamically adjusts to the characteristics of deep learning data. AFP requires no changes to the model topology, requires no additional training, and applies to all layers of DNN models. We evaluate AFP on a spectrum of representative models in computer vision and NLP, and show that our technique enables ultra-low precision inference of deep learning models while providing accuracy comparable to full precision inference. By dynamically adjusting to ML data, AFP increases memory density by 1.6x, 1.6x, and 3.2x and compute density by 4x, 1.3x, and 12x when compared to BFP, BFloat16, and FP32.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- BiE: Bi-Exponent Block Floating-Point for Large Language Models QuantizationLancheng Zou, Wenqian Zhao, Shuo Yin, Chen Bai 等ICML 2024 · 被引用 13 次
- MX+: Pushing the Limits of Microscaling Formats for Efficient Large Language Model ServingJungi Lee, Junyong Park, Soohyun Cha, Jaehoon Cho 等MICRO 2025 · 被引用 7 次
- BBAL: A Bidirectional Block Floating Point-Based Quantisation Accelerator for Large Language ModelsXiaomeng Han, Yuan Cheng, Jing Wang, Junyang Lu 等DAC 2025 · 被引用 5 次
- Pushing the Limits of BFP on Narrow Precision LLM InferenceHui Wang, Yuan Cheng, Xiaomeng Han, Zhengpeng Zhao 等AAAI 2025 · 被引用 1 次
- Effective Interplay between Sparsity and Quantization: From Theory to PracticeSimla Burcu Harma, Ayan Chakraborty, Elizaveta Kostenok, Danila Mishin 等ICLR 2025
它引用的顶会 Paper7
- Ultra-Low Precision 4-bit Training of Deep Neural NetworksXiao Sun, Naigang Wang, Chia-Yu Chen, Jiamin Ni 等NeurIPS 2020 · 被引用 227 次
- Pushing the Limits of Narrow Precision Inferencing at Cloud Scale with Microsoft Floating PointBita Darvish Rouhani, Daniel Lo, Ritchie Zhao, Ming Liu 等NeurIPS 2020 · 被引用 153 次
- Mix and Match: A Novel FPGA-Centric Deep Neural Network Quantization FrameworkSung-En Chang, Yanyu Li, Mengshu Sun, Runbin Shi 等HPCA 2021 · 被引用 125 次
- Think Fast: A Tensor Streaming Processor (TSP) for Accelerating Deep Learning WorkloadsDennis Abts, Jonathan Ross, Jonathan Sparling, Mark Wong-VanHaren 等ISCA 2020 · 被引用 91 次
- RaPiD: AI Accelerator for Ultra-low Precision Training and InferenceSwagath Venkataramani, Vijayalakshmi Srinivasan, Wei Wang, Sanchari Sen 等ISCA 2021 · 被引用 76 次
相关 Paper
- Improving Neural Network Efficiency via Post-training Quantization with Adaptive Floating-PointFangxin Liu, Wenbo Zhao, Zhezhi He, Yanzhi Wang 等ICCV 2021 · 被引用 69 次
- Any-Precision Deep Neural NetworksHaichao Yu, Haoxiang Li, Humphrey Shi, Thomas S. Huang 等AAAI 2021 · 被引用 79 次
- Fixed-Point Back-Propagation TrainingXishan Zhang, Shaoli Liu, Rui Zhang, Chang Liu 等CVPR 2020
- Algorithm-Hardware Co-Design of Adaptive Floating-Point Encodings for Resilient Deep Learning InferenceThierry Tambe, En-Yu Yang, Zishen Wan, Yuntian Deng 等DAC 2020 · 被引用 67 次
- Drift: Leveraging Distribution-based Dynamic Precision Quantization for Efficient Deep Neural Network AccelerationLian Liu, Zhaohui Xu, Yintao He, Ying Wang 等DAC 2024 · 被引用 5 次
