RecFlow: Unlocking GPU Efficiency for DLRM Inference via Fine-Grained Parallelism and Incremental Batching
Siheng Pan, Shaolong Li, Minwei Zhang, Shuxi Guo, Haifeng Sun, Qi Qi, Zirui Zhuang, Xiang Yang, Chunyang Jiang, Jianxin Liao, Jing Wang
摘要
Industrial Deep Learning Recommendation Models (DLRMs) comprise memory-intensive embedding operations and compute-intensive DNN layers, often resulting in suboptimal GPU resource utilization under high-throughput inference workloads. However, the memory demands of DNN layers vary considerably across different workloads and execution phases, often leading to unpredictable interference, which makes it challenging to efficiently co-execute embedding and DNN operators.This paper presents RecFlow, a high-performance DLRM serving framework that leverages intra-SM parallelism to co-run embedding and DNN computations through fine-grained resource coordination. RecFlow profiles the workload characteristics of each DNN phase and applies adaptive parallel strategies to sustain high memory bandwidth utilization while minimizing interoperator interference. To further enable parallelism between the structurally dependent embedding and top-DNN stages, RecFlow introduces an incremental batching mechanism that overlaps their execution using newly arrived requests, thereby enabling inter-batch parallelism without incurring additional queuing latency. Extensive evaluations on real-world production workloads demonstrate that RecFlow improves serving throughput by up to 1.13 × over state-of-the-art DLRM inference systems while reducing latency in high-throughput serving scenarios.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Optimizing CPU Performance for Recommendation Systems At-ScaleRishabh Jain, Scott Cheng, Vishwas Kalagi, Vrushabh Sanghavi 等ISCA 2023 · 被引用 25 次
- UpDLRM: Accelerating Personalized Recommendation using Real-World PIM ArchitectureSitian Chen, Haobin Tan, Amelie Chi Zhou, Yusen Li 等DAC 2024 · 被引用 9 次
- Pushing the Performance Envelope of DNN-based Recommendation Systems Inference on GPUsRishabh Jain, Vivek M. Bhasi, Adwait Jog, Anand Sivasubramaniam 等MICRO 2024 · 被引用 5 次
- RecPipe: Co-designing Models and Hardware to Jointly Optimize Recommendation Quality and PerformanceUdit Gupta, Samuel Hsia, Jeff Zhang, Mark Wilkening 等MICRO 2021 · 被引用 31 次
- RecNMP: Accelerating Personalized Recommendation with Near-Memory ProcessingLiu Ke, Udit Gupta, Benjamin Youngjae Cho, David Brooks 等ISCA 2020 · 被引用 235 次
