RAP: Resource-aware Automated GPU Sharing for Multi-GPU Recommendation Model Training and Input Preprocessing
Zheng Wang, Yuke Wang, Jiaqi Deng, Da Zheng, Ang Li, Yufei Ding
Abstract
Ensuring high-quality recommendations for newly onboarded users requires the continuous retraining of Deep Learning Recommendation Models (DLRMs) with freshly generated data. To serve the online DLRM retraining, existing solutions use hundreds of CPU computing nodes designated for input preprocessing, causing significant power consumption that surpasses even the power usage of GPU trainers.
To this end, we propose RAP, an end-to-end DLRM training framework that supports Resource-aware Automated GPU sharing for DLRM input Preprocessing and Training. The core idea of RAP is to accurately capture the remaining GPU computing resources during DLRM training for input preprocessing, achieving superior training efficiency without requiring additional resources. Specifically, RAP utilizes a co-running cost model to efficiently assess the costs of various input preprocessing operations, and it implements a resource-aware horizontal fusion technique that adaptively merges smaller kernels according to GPU availability, circumventing any interference with DLRM training. In addition, RAP leverages a heuristic searching algorithm that jointly optimizes both the input preprocessing graph mapping and the co-running schedule to maximize the end-to-end DLRM training throughput. The comprehensive evaluation shows that RAP achieves 2.09× speedup on average over the sequential GPU-based DLRM input preprocessing baseline. In addition, the end-to-end training throughput of RAP is only
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2d0f21ad-2e7c-4b8c-a357-3c1a6bf20c2eCited by top-tier papers8
- GPU-Disaggregated Serving for Deep Learning Recommendation Models at ScaleLingyun Yang, Yongchen Wang, Yinghao Yu, Qizhen Weng et al.NSDI 2025 · 22 citations
- OPER: Optimality-Guided Embedding Table Parallelization for Large-scale Recommendation ModelZheng Wang, Yuke Wang, Boyuan Feng, Guyue Huang et al.USENIX ATC 2024 · 7 citations
- Machine Learning-Guided Memory Optimization for DLRM Inference on Tiered MemoryJie Ren, Bin Ma, Shuangyan Yang, Benjamin Francis et al.HPCA 2025 · 6 citations
- RecFlex: Enabling Feature Heterogeneity-Aware Optimization for Deep Recommendation Models with Flexible SchedulesZaifeng Pan, Zhen Zheng, Feng Zhang, Bing Xie et al.SC 2024 · 2 citations
- Bat: Efficient Generative Recommender Serving with Bipartite AttentionJie Sun, Shaohang Wang, Zimo Zhang, Zhengyu Liu et al.ASPLOS 2026 · 1 citation
Builds on12
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 852 citations
- Microsecond-scale Preemption for Concurrent GPU-accelerated DNN InferencesMingcong Han, Hanze Zhang, Rong Chen, Haibo ChenOSDI 2022 · 153 citations
- Analyzing and Mitigating Data Stalls in DNN TrainingJayashree Mohan, Amar Phanishayee, Ashish Raniwala, Vijay ChidambaramVLDB 2021 · 142 citations
- Accelerating Recommendation System Training by Leveraging Popular ChoicesMuhammad Adnan, Yassaman Ebrahimzadeh Maboud, Divya Mahajan, Prashant J. NairVLDB 2022 · 70 citations
- RecShard: statistical feature-based memory optimization for industry-scale neural recommendationGeet Sethi, Bilge Acun, Niket Agarwal, Christos Kozyrakis et al.ASPLOS 2022 · 65 citations
Related papers
- RecFlow: Unlocking GPU Efficiency for DLRM Inference via Fine-Grained Parallelism and Incremental BatchingSiheng Pan, Shaolong Li, Minwei Zhang, Shuxi Guo et al.INFOCOM 2026
- EL-Rec: Efficient Large-Scale Recommendation Model Training via Tensor-Train Embedding TableZheng Wang, Yuke Wang, Boyuan Feng, Dheevatsa Mudigere et al.SC 2022 · 15 citations
- FusedRec: Fused Embedding Communication for Distributed Recommendation Training on GPUsXuanteng Huang, Fan Li, Riyang Hu, Jianchang Zhang et al.AAAI 2026 · 1 citation
- Accelerating Neural Recommendation Training with Embedding SchedulingChaoliang Zeng, Xudong Liao, Xiaodian Cheng, Han Tian et al.NSDI 2024 · 16 citations
- HypeReca: Distributed Heterogeneous In-Memory Embedding Database for Training Recommender ModelsJiaao He, Shengqi Chen, Kezhao Huang, Jidong ZhaiUSENIX ATC 2025 · 2 citations
