EagleRec: Edge-Scale Recommendation System Acceleration with Inter-Stage Parallelism Optimization on GPUs
Yongbo Yu, Fuxun Yu, Xiang Sheng, Chenchen Liu, Xiang Chen
摘要
Recommendation systems suggest items to users by predicting their preferences based on historical data. The industry traditionally handles large-scale recommendation requests by scaling the number of devices without much concern for a single device’s performance. However, there is a trend for recommendation systems to gradually move from a centralized service to an edge device. The edge-scale recommendation systems have distinct features that are different from traditional large-scale deployments, which poses different challenges to the acceleration of the recommendation system. In this paper, we focus on the edge-scale recommendation system and propose an inter-stage parallelism optimization method deployed on a single GPU. Experiments show that our framework could improve recommendation system throughput by 1.89× 2.2× for different datasets on the GPU.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- RecPipe: Co-designing Models and Hardware to Jointly Optimize Recommendation Quality and PerformanceUdit Gupta, Samuel Hsia, Jeff Zhang, Mark Wilkening 等MICRO 2021 · 被引用 31 次
- Fleche: an efficient GPU embedding cache for personalized recommendationsMinhui Xie, Youyou Lu, Jiazhen Lin, Qing Wang 等EuroSys 2022 · 被引用 24 次
- InfScaler: Enabling Efficient ML Inference Serving on Multi-Accelerator Edge Devices via Asymmetric Auto-ScalingBorui Li, Tiange Xia, Shuai Wang, Shuai WangDAC 2025 · 被引用 2 次
- RecFlow: Unlocking GPU Efficiency for DLRM Inference via Fine-Grained Parallelism and Incremental BatchingSiheng Pan, Shaolong Li, Minwei Zhang, Shuxi Guo 等INFOCOM 2026
- DeepRecSys: A System for Optimizing End-To-End At-Scale Neural Recommendation InferenceUdit Gupta, Samuel Hsia, Vikram Saraph, Xiaodong Wang 等ISCA 2020 · 被引用 149 次
