Campo: Cost-Aware Performance Optimization for Mixed-Precision Neural Network Training
Xin He, Jianhua Sun, Hao Chen, Dong Li
Abstract
Mixed precision training uses a mixture of full and lower precisions for neural network (NN) training. Applying mixed precision must cast tensors in NN from float32 (FP32) to float16 (FP16) or vice versa. The existing strategy greedily applies FP16 to performance-critical operations without quantifying and considering the casting cost. However, we reveal that the casting cost can take more than 21% of NN operation execution time, and in some cases surpasses the performance benefit of using low precision. In this paper, we introduce Campo, a tool that improves performance of mixed-precision NN training with the awareness of casting costs. Campo is built upon performance modeling that predicts the casting cost and operation performance with low precision, and introduces a cost-aware graph rewriting strategy. Campo is user-transparent, and enables high performance NN training using mixed precision without training accuracy loss. Evaluating Campo with six NN models, we show that compared to TensorFlow using TF_AMP (a state-of-the-art performance optimizer for mixed precision training from Nvidia), Campo improves training throughput by 20.8% on average (up to 24.5%) on RTX 2080 Ti GPU and by 20.9% on average (up to 23.4%) on V100 GPU, without training accuracy loss. Because of using the cost-aware mixed precision training, Campo also improves energy efficiency by 21.4% on average (up to 24.2%), compared to TensorFlow using TF_AMP.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d3c6cfdf-4719-45b7-af26-71060e2f36d4Cited by top-tier papers3
- Why Low-Precision Transformer Training Fails: An Analysis on Flash AttentionHaiquan Qiu, Quanming YaoICLR 2026 · 15 citations
- AdaCheck: An Adaptive Checkpointing System for Efficient LLM Training with Redundancy UtilizationWeijie Liu, Shengwei Li, Zhiquan Lai, Keshi Ge et al.FAST 2026 · 3 citations
- WATOS: Efficient LLM Training Strategies and Architecture Co-Exploration for Wafer-Scale ChipHuizheng Wang, Zichuan Wang, Hongbin Wang, Jingxiang Hou et al.HPCA 2026 · 2 citations
Builds on2
- Capuchin: Tensor-based GPU Memory Management for Deep LearningXuan Peng, Xuanhua Shi, Hulin Dai, Hai Jin et al.ASPLOS 2020 · 143 citations
- AutoTM: Automatic Tensor Movement in Heterogeneous Memory Systems using Integer Linear ProgrammingMark Hildebrand, Jawad Khan, Sanjeev Trika, Jason Lowe-Power et al.ASPLOS 2020 · 70 citations
Related papers
- Multi-Precision Policy Enforced Training (MuPPET) : A Precision-Switching Strategy for Quantised Fixed-Point Training of CNNsAditya Rajagopal, Diederik Adriaan Vink, Stylianos I. Venieris, Christos-Savvas BouganisICML 2020 · 17 citations
- Predicting Performance and Accuracy of Mixed-Precision Programs for Precision TuningYutong Wang, Cindy Rubio-GonzálezICSE 2024 · 7 citations
- Efficient Quantized Sparse Matrix Operations on Tensor CoresShigang Li, Kazuki Osawa, Torsten HoeflerSC 2022 · 27 citations
- AMPA: Adaptive Mixed Precision Allocation for Low-Bit Integer TrainingLi Ding, Wen Fei, Yuyang Huang, Shuangrui Ding et al.ICML 2024 · 5 citations
- Drift: Leveraging Distribution-based Dynamic Precision Quantization for Efficient Deep Neural Network AccelerationLian Liu, Zhaohui Xu, Yintao He, Ying Wang et al.DAC 2024 · 5 citations
