USENIX ATC2022顶会
Campo: Cost-Aware Performance Optimization for Mixed-Precision Neural Network Training
Xin He, Jianhua Sun, Hao Chen, Dong Li
摘要
Mixed precision training uses a mixture of full and lower precisions for neural network (NN) training. Applying mixed precision must cast tensors in NN from float32 (FP32) to float16 (FP16) or vice versa. The existing strategy greedily applies FP16 to performance-critical operations without quantifying and considering the casting cost. However, we reveal that the casting cost can take more than 21% of NN operation execution time, and in some cases surpasses the performance benefit of using low precision. In this paper, we introduce Campo, a tool that improves performance of mixed-precision NN training with the awareness of casting costs. Campo is built upon performance modeling that predicts the casting cost and operation performance with low precision, and introduces a cost-aware graph rewriting strategy. Campo is user-transparent, and enables high performance NN training using mixed precision without training accuracy loss. Evaluating Campo with six NN models, we show that compared to TensorFlow using TF_AMP (a state-of-the-art performance optimizer for mixed precision training from Nvidia), Campo improves training throughput by 20.8% on average (up to 24.5%) on RTX 2080 Ti GPU and by 20.9% on average (up to 23.4%) on V100 GPU, without training accuracy loss. Because of using the cost-aware mixed precision training, Campo also improves energy efficiency by 21.4% on average (up to 24.2%), compared to TensorFlow using TF_AMP.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Why Low-Precision Transformer Training Fails: An Analysis on Flash AttentionHaiquan Qiu, Quanming YaoICLR 2026 · 被引用 15 次
- AdaCheck: An Adaptive Checkpointing System for Efficient LLM Training with Redundancy UtilizationWeijie Liu, Shengwei Li, Zhiquan Lai, Keshi Ge 等FAST 2026 · 被引用 3 次
- WATOS: Efficient LLM Training Strategies and Architecture Co-Exploration for Wafer-Scale ChipHuizheng Wang, Zichuan Wang, Hongbin Wang, Jingxiang Hou 等HPCA 2026 · 被引用 2 次
它引用的顶会 Paper2
- Capuchin: Tensor-based GPU Memory Management for Deep LearningXuan Peng, Xuanhua Shi, Hulin Dai, Hai Jin 等ASPLOS 2020 · 被引用 143 次
- AutoTM: Automatic Tensor Movement in Heterogeneous Memory Systems using Integer Linear ProgrammingMark Hildebrand, Jawad Khan, Sanjeev Trika, Jason Lowe-Power 等ASPLOS 2020 · 被引用 70 次
相关 Paper
- Multi-Precision Policy Enforced Training (MuPPET) : A Precision-Switching Strategy for Quantised Fixed-Point Training of CNNsAditya Rajagopal, Diederik Adriaan Vink, Stylianos I. Venieris, Christos-Savvas BouganisICML 2020 · 被引用 17 次
- Predicting Performance and Accuracy of Mixed-Precision Programs for Precision TuningYutong Wang, Cindy Rubio-GonzálezICSE 2024 · 被引用 7 次
- Efficient Quantized Sparse Matrix Operations on Tensor CoresShigang Li, Kazuki Osawa, Torsten HoeflerSC 2022 · 被引用 27 次
- AMPA: Adaptive Mixed Precision Allocation for Low-Bit Integer TrainingLi Ding, Wen Fei, Yuyang Huang, Shuangrui Ding 等ICML 2024 · 被引用 5 次
- Drift: Leveraging Distribution-based Dynamic Precision Quantization for Efficient Deep Neural Network AccelerationLian Liu, Zhaohui Xu, Yintao He, Ying Wang 等DAC 2024 · 被引用 5 次
