FPRaker: A Processing Element For Accelerating Neural Network Training
Omar Mohamed Awad, Mostafa Mahmoud, Isak Edo, Ali Hadi Zadeh, Ciaran Bannon, Anand Jayarajan, Gennady Pekhimenko, Andreas Moshovos
摘要
We present FPRaker, a processing element for composing training accelerators. FPRaker processes several floating-point multiply-accumulation operations concurrently and accumulates their result into a higher precision accumulator. FPRaker boosts performance and energy efficiency during training by taking advantage of the values that naturally appear during training. It processes the significand of the operands of each multiply-accumulate as a series of signed powers of two. The conversion to this form is done on-the-fly. This exposes ineffectual work that can be skipped: values when encoded have few terms and some of them can be discarded as they would fall outside the range of the accumulator given the limited precision of floating-point. FPRaker also takes advantage of spatial correlation in values across channels and uses delta-encoding off-chip to reduce memory footprint and bandwidth. We demonstrate that FPRaker can be used to compose an accelerator for training and that it can improve performance and energy efficiency compared to using optimized bit-parallel floating-point units under iso-compute area constraints. We also demonstrate that FPRaker delivers additional benefits when training incorporates pruning and quantization. Finally, we show that FPRaker naturally amplifies performance with training methods that use a different precision per layer.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- FAST: DNN Training Under Variable Precision Block Floating Point with Stochastic RoundingSai Qian Zhang, Bradley McDanel, H. T. KungHPCA 2022 · 被引用 68 次
- TensorDash: Exploiting Sparsity to Accelerate Deep Neural Network TrainingMostafa Mahmoud, Isak Edo, Ali Hadi Zadeh, Omar Mohamed Awad 等MICRO 2020 · 被引用 78 次
- A2Q: Accumulator-Aware Quantization with Guaranteed Overflow AvoidanceIan Colbert, Alessandro Pappalardo, Jakoba Petri-KoenigICCV 2023 · 被引用 19 次
- Cambricon-Q: A Hybrid Architecture for Efficient TrainingYongwei Zhao, Chang Liu, Zidong Du, Qi Guo 等ISCA 2021 · 被引用 28 次
- Procrustes: a Dataflow and Accelerator for Sparse Deep Neural Network TrainingDingqing Yang, Amin Ghasemazar, Xiaowei Ren, Maximilian Golub 等MICRO 2020 · 被引用 63 次
