RecFlow: Unlocking GPU Efficiency for DLRM Inference via Fine-Grained Parallelism and Incremental Batching
Siheng Pan, Shaolong Li, Minwei Zhang, Shuxi Guo, Haifeng Sun, Qi Qi, Zirui Zhuang, Xiang Yang, Chunyang Jiang, Jianxin Liao, Jing Wang
Abstract
Industrial Deep Learning Recommendation Models (DLRMs) comprise memory-intensive embedding operations and compute-intensive DNN layers, often resulting in suboptimal GPU resource utilization under high-throughput inference workloads. However, the memory demands of DNN layers vary considerably across different workloads and execution phases, often leading to unpredictable interference, which makes it challenging to efficiently co-execute embedding and DNN operators.This paper presents RecFlow, a high-performance DLRM serving framework that leverages intra-SM parallelism to co-run embedding and DNN computations through fine-grained resource coordination. RecFlow profiles the workload characteristics of each DNN phase and applies adaptive parallel strategies to sustain high memory bandwidth utilization while minimizing interoperator interference. To further enable parallelism between the structurally dependent embedding and top-DNN stages, RecFlow introduces an incremental batching mechanism that overlaps their execution using newly arrived requests, thereby enabling inter-batch parallelism without incurring additional queuing latency. Extensive evaluations on real-world production workloads demonstrate that RecFlow improves serving throughput by up to 1.13 × over state-of-the-art DLRM inference systems while reducing latency in high-throughput serving scenarios.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get d48a4078-a124-488d-ac26-357ea2db704bRelated papers
- Optimizing CPU Performance for Recommendation Systems At-ScaleRishabh Jain, Scott Cheng, Vishwas Kalagi, Vrushabh Sanghavi et al.ISCA 2023 · 25 citations
- UpDLRM: Accelerating Personalized Recommendation using Real-World PIM ArchitectureSitian Chen, Haobin Tan, Amelie Chi Zhou, Yusen Li et al.DAC 2024 · 9 citations
- Pushing the Performance Envelope of DNN-based Recommendation Systems Inference on GPUsRishabh Jain, Vivek M. Bhasi, Adwait Jog, Anand Sivasubramaniam et al.MICRO 2024 · 5 citations
- RecPipe: Co-designing Models and Hardware to Jointly Optimize Recommendation Quality and PerformanceUdit Gupta, Samuel Hsia, Jeff Zhang, Mark Wilkening et al.MICRO 2021 · 31 citations
- RecNMP: Accelerating Personalized Recommendation with Near-Memory ProcessingLiu Ke, Udit Gupta, Benjamin Youngjae Cho, David Brooks et al.ISCA 2020 · 235 citations
