Fleche: an efficient GPU embedding cache for personalized recommendations
Minhui Xie, Youyou Lu, Jiazhen Lin, Qing Wang, Jian Gao, Kai Ren, Jiwu Shu
摘要
Deep learning based models have dominated current production recommendation systems. However, the gap between CPU-side DRAM data accessing and GPU processing still impedes their inference performance. GPU-resident cache can bridge this gap, but we find that existing systems leave the benefits to cache the embedding table, a huge sparse structure, on GPU unexploited. In this paper, we present Fleche, a holistic cache scheme with detailed designs for efficient GPU-resident embedding caching. Fleche (1) uses one cache backend for all embedding tables to improve the total cache utilization, and (2) merges small kernel calls into one unitary call to reduce the overhead of kernel maintenance (e.g., kernel launching and synchronizing). Furthermore, we carefully design the cache query workflow for finer-grain parallelism. Evaluations with real-world datasets show that compared with the prior art, Fleche significantly improves the throughput of embedding layer by 2.0 -- 5.4×, and gets up to 2.4× speedup of end-to-end inference throughput.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- AdaEmbed: Adaptive Embedding for Large-Scale Recommendation ModelsFan Lai, Wei Zhang, Rui Liu, William Tsai 等OSDI 2023 · 被引用 23 次
- GPU-Disaggregated Serving for Deep Learning Recommendation Models at ScaleLingyun Yang, Yongchen Wang, Yinghao Yu, Qizhen Weng 等NSDI 2025 · 被引用 22 次
- Fast State Restoration in LLM Serving with HCacheShiwei Gao, Youmin Chen, Jiwu ShuEuroSys 2025 · 被引用 22 次
- UGACHE: A Unified GPU Cache for Embedding-based Deep LearningXiaoniu Song, Yiwen Zhang, Rong Chen, Haibo ChenSOSP 2023 · 被引用 20 次
- Optimizing Distributed ML Communication with Fused Computation-Collective OperationsKishore Punniyamurthy, Khaled Hamidouche, Bradford M. BeckmannSC 2024 · 被引用 11 次
它引用的顶会 Paper7
- RecNMP: Accelerating Personalized Recommendation with Near-Memory ProcessingLiu Ke, Udit Gupta, Benjamin Youngjae Cho, David Brooks 等ISCA 2020 · 被引用 235 次
- Centaur: A Chiplet-based, Hybrid Sparse-Dense Accelerator for Personalized RecommendationsRanggi Hwang, Taehun Kim, Youngeun Kwon, Minsoo RhuISCA 2020 · 被引用 94 次
- FAFNIR: Accelerating Sparse Gathering by Using Efficient Near-Memory Intelligent ReductionBahar Asgari, Ramyad Hadidi, Jiashen Cao, Da Eun Shim 等HPCA 2021 · 被引用 87 次
- SPACE: Locality-Aware Processing in Heterogeneous Memory for Personalized RecommendationsHongju Kal, Seokmin Lee, Gun Ko, Won Woo RoISCA 2021 · 被引用 43 次
- Kraken: memory-efficient continual learning for large-scale real-time recommendationsMinhui Xie, Kai Ren, Youyou Lu, Guangxu Yang 等SC 2020 · 被引用 38 次
相关 Paper
- Accelerating Recommendation System Training by Leveraging Popular ChoicesMuhammad Adnan, Yassaman Ebrahimzadeh Maboud, Divya Mahajan, Prashant J. NairVLDB 2022 · 被引用 70 次
- Training personalized recommendation systems from (GPU) scratch: look forward not backwardsYoungeun Kwon, Minsoo RhuISCA 2022 · 被引用 24 次
- Tensor Casting: Co-Designing Algorithm-Architecture for Personalized Recommendation TrainingYoungeun Kwon, Yunjae Lee, Minsoo RhuHPCA 2021 · 被引用 40 次
- RecPipe: Co-designing Models and Hardware to Jointly Optimize Recommendation Quality and PerformanceUdit Gupta, Samuel Hsia, Jeff Zhang, Mark Wilkening 等MICRO 2021 · 被引用 31 次
- Bagpipe: Accelerating Deep Recommendation Model TrainingSaurabh Agarwal, Chengpo Yan, Ziyi Zhang, Shivaram VenkataramanSOSP 2023 · 被引用 17 次
