CowClip: Reducing CTR Prediction Model Training Time from 12 Hours to 10 Minutes on 1 GPU
Zangwei Zheng, Pengtai Xu, Xuan Zou, Da Tang, Zhen Li, Chenguang Xi, Peng Wu, Leqi Zou, Yijie Zhu, Ming Chen, Xiangzhuo Ding, Fuzhao Xue
Abstract
The click-through rate (CTR) prediction task is to predict whether a user will click on the recommended item. As mind-boggling amounts of data are produced online daily, accelerating CTR prediction model training is critical to ensuring an up-to-date model and reducing the training cost. One approach to increase the training speed is to apply large batch training. However, as shown in computer vision and natural language processing tasks, training with a large batch easily suffers from the loss of accuracy. Our experiments show that previous scaling rules fail in the training of CTR prediction neural networks. To tackle this problem, we first theoretically show that different frequencies of ids make it challenging to scale hyperparameters when scaling the batch size. To stabilize the training process in a large batch size setting, we develop the adaptive Column-wise Clipping (CowClip). It enables an easy and effective scaling rule for the embeddings, which keeps the learning rate unchanged and scales the L2 loss. We conduct extensive experiments with four CTR prediction networks on two real-world datasets and successfully scaled 128 times the original batch size without accuracy loss. In particular, for CTR prediction model DeepFM training on the Criteo dataset, our optimization framework enlarges the batch size from 1K to 128K with over 0.1% AUC improvement and reduces training time from 12 hours to 10 minutes on a single V100 GPU. Our code locates at github.com/bytedance/LargeBatchCTR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2a0c208b-b367-49d0-aa1e-2e7905bc58fbCited by top-tier papers1
Ask how each one uses itBuilds on8
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu et al.ICLR 2020 · 1,170 citations
- DCN V2: Improved Deep & Cross Network and Practical Lessons for Web-scale Learning to Rank SystemsRuoxi Wang, Rakesh Shivanna, Derek Zhiyuan Cheng, Sagar Jain et al.WWW 2021 · 793 citations
- High-Performance Large-Scale Image Recognition Without NormalizationAndy Brock, Soham De, Samuel L. Smith, Karen SimonyanICML 2021 · 613 citations
- Why Gradient Clipping Accelerates Training: A Theoretical Justification for AdaptivityJingzhao Zhang, Tianxing He, Suvrit Sra, Ali JadbabaieICLR 2020 · 598 citations
- Accelerating Recommendation System Training by Leveraging Popular ChoicesMuhammad Adnan, Yassaman Ebrahimzadeh Maboud, Divya Mahajan, Prashant J. NairVLDB 2022 · 70 citations
Related papers
- Adaptive Low-Precision Training for Embeddings in Click-Through Rate PredictionShiwei Li, Huifeng Guo, Lu Hou, Wei Zhang et al.AAAI 2023 · 27 citations
- Agile and Accurate CTR Prediction Model Training for Massive-Scale Online Advertising SystemsZhiqiang Xu, Dong Li, Weijie Zhao, Xing Shen et al.SIGMOD 2021 · 38 citations
- ScaleFreeCTR: MixCache-based Distributed Training System for CTR Models with Huge Embedding TableHuifeng Guo, Wei Guo, Yong Gao, Ruiming Tang et al.SIGIR 2021 · 24 citations
- Dual Graph enhanced Embedding Neural Network for CTR PredictionWei Guo, Rong Su, Renhao Tan, Huifeng Guo et al.KDD 2021 · 72 citations
- Looking at CTR Prediction Again: Is Attention All You Need?Yuan Cheng, Yanbo XueSIGIR 2021 · 18 citations
