Lune

INFOCOM2026顶会

RecFlow: Unlocking GPU Efficiency for DLRM Inference via Fine-Grained Parallelism and Incremental Batching

Siheng Pan, Shaolong Li, Minwei Zhang, Shuxi Guo, Haifeng Sun, Qi Qi, Zirui Zhuang, Xiang Yang, Chunyang Jiang, Jianxin Liao, Jing Wang

2026年份

摘要

Industrial Deep Learning Recommendation Models (DLRMs) comprise memory-intensive embedding operations and compute-intensive DNN layers, often resulting in suboptimal GPU resource utilization under high-throughput inference workloads. However, the memory demands of DNN layers vary considerably across different workloads and execution phases, often leading to unpredictable interference, which makes it challenging to efficiently co-execute embedding and DNN operators.This paper presents RecFlow, a high-performance DLRM serving framework that leverages intra-SM parallelism to co-run embedding and DNN computations through fine-grained resource coordination. RecFlow profiles the workload characteristics of each DNN phase and applies adaptive parallel strategies to sustain high memory bandwidth utilization while minimizing interoperator interference. To further enable parallelism between the structurally dependent embedding and top-DNN stages, RecFlow introduces an incremental batching mechanism that overlaps their execution using newly arrived requests, thereby enabling inter-batch parallelism without incurring additional queuing latency. Extensive evaluations on real-world production workloads demonstrate that RecFlow improves serving throughput by up to 1.13 × over state-of-the-art DLRM inference systems while reducing latency in high-throughput serving scenarios.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

lune papers get d48a4078-a124-488d-ac26-357ea2db704b

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖