Faster Neural Network Training with Approximate Tensor Operations
Menachem Adelman, Kfir Y. Levy, Ido Hakimi, Mark Silberstein
摘要
We propose a novel technique for faster deep neural network training which systematically applies sample-based approximation to the constituent tensor operations, i.e., matrix multiplications and convolutions. We introduce new sampling techniques, study their theoretical properties, and prove that they provide the same convergence guarantees when applied to SGD training. We apply approximate tensor operations to single and multi-node training of MLP and CNN networks on MNIST, CIFAR-10 and ImageNet datasets. We demonstrate up to 66% reduction in the amount of computations and communication, and up to 1.37x faster training time while maintaining negligible or no impact on the final test accuracy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Fine-Tuning Language Models with Just Forward PassesSadhika Malladi, Tianyu Gao, Eshaan Nichani, Alex Damian 等NeurIPS 2023 · 被引用 495 次
- Training Transformers with 4-bit IntegersHaocheng Xi, Changhao Li, Jianfei Chen, Jun ZhuNeurIPS 2023 · 被引用 96 次
- L2ight: Enabling On-Chip Learning for Optical Neural Networks via Efficient in-situ Subspace OptimizationJiaqi Gu, Hanqing Zhu, Chenghao Feng, Zixuan Jiang 等NeurIPS 2021 · 被引用 41 次
- Randomized Automatic DifferentiationDeniz Oktay, Nick McGreivy, Joshua Aduol, Alex Beatson 等ICLR 2021 · 被引用 31 次
- RSC: Accelerate Graph Neural Networks Training via Randomized Sparse ComputationsZirui Liu, Shengyuan Chen, Kaixiong Zhou, Daochen Zha 等ICML 2023 · 被引用 24 次
它引用的顶会 Paper4
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
- Accelerated Sparse Neural Training: A Provable and Efficient Method to Find N: M Transposable MasksItay Hubara, Brian Chmiel, Moshe Island, Ron Banner 等NeurIPS 2021 · 被引用 148 次
- Rethinking Attention with PerformersKrzysztof Marcin Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song 等ICLR 2021 · 被引用 122 次
- ReSprop: Reuse Sparsified BackpropagationNegar Goli, Tor M. AamodtCVPR 2020
相关 Paper
- A Layer-Wise Natural Gradient Optimizer for Training Deep Neural NetworksXiaolei Liu, Shaoshuai Li, Kaixin Gao, Binfeng WangNeurIPS 2024 · 被引用 2 次
- Control Variate Approximation for DNN AcceleratorsGeorgios Zervakis, Ourania Spantidi, Iraklis Anagnostopoulos, Hussam Amrouch 等DAC 2021 · 被引用 32 次
- FAST: DNN Training Under Variable Precision Block Floating Point with Stochastic RoundingSai Qian Zhang, Bradley McDanel, H. T. KungHPCA 2022 · 被引用 68 次
- Stochastic Weight Averaging in Parallel: Large-Batch Training That Generalizes WellVipul Gupta, Santiago Akle Serrano, Dennis DeCosteICLR 2020 · 被引用 78 次
- Prediction Confidence based Low Complexity Gradient Computation for Accelerating DNN TrainingDongyeob Shin, Geonho Kim, Joongho Jo, Jongsun ParkDAC 2020 · 被引用 14 次
