RECom: A Compiler Approach to Accelerating Recommendation Model Inference with Massive Embedding Columns
Zaifeng Pan, Zhen Zheng, Feng Zhang, Ruofan Wu, Hao Liang, Dalin Wang, Xiafei Qiu, Junjie Bai, Wei Lin, Xiaoyong Du
摘要
Embedding columns are important for deep recommendation models to achieve high accuracy, but they can be very time-consuming during inference. Machine learning (ML) compilers are used broadly in real businesses to optimize ML models automatically. Unfortunately, no existing work uses compilers to automatically accelerate the heavy embedding column computations during recommendation model inferences. To fill this gap, we propose RECom, the first ML compiler that aims at optimizing the massive embedding columns in recommendation models on the GPU. RECom addresses three major challenges. First, generating an efficient schedule on the GPU for the massive operators within embedding columns is difficult. Existing solutions usually lead to numerous small kernels and also lack inter-subgraph parallelism. We adopt a novel codegen strategy that fuses massive embedding columns into a single kernel and maps each column into a separate thread block on the GPU. Second, the complex shape computations under dynamic shape scenarios impede further graph optimizations. We develop a symbolic expression-based module to reconstruct all shape computations. Third, ML frameworks inevitably introduce redundant computations due to robustness considerations. We develop a subgraph optimization module that performs graph-level simplifications based on the entire embedding column context. Experiments on both in-house and open-source models show that RECom can achieve 6.61X and 1.91X over state-of-the-art baselines in terms of end-to-end inference latency and throughput, respectively. RECom's source code is publicly available at https://github.com/AlibabaResearch/recom.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper6
- MonoNN: Enabling a New Monolithic Optimization Space for Neural Network Inference Tasks on Modern GPU-Centric ArchitecturesDonglin Zhuang, Zhen Zheng, Haojun Xia, Xiafei Qiu 等OSDI 2024 · 被引用 11 次
- Machine Learning-Guided Memory Optimization for DLRM Inference on Tiered MemoryJie Ren, Bin Ma, Shuangyan Yang, Benjamin Francis 等HPCA 2025 · 被引用 6 次
- RecFlex: Enabling Feature Heterogeneity-Aware Optimization for Deep Recommendation Models with Flexible SchedulesZaifeng Pan, Zhen Zheng, Feng Zhang, Bing Xie 等SC 2024 · 被引用 2 次
- FusedRec: Fused Embedding Communication for Distributed Recommendation Training on GPUsXuanteng Huang, Fan Li, Riyang Hu, Jianchang Zhang 等AAAI 2026 · 被引用 1 次
- Bat: Efficient Generative Recommender Serving with Bipartite AttentionJie Sun, Shaohang Wang, Zimo Zhang, Zhengyu Liu 等ASPLOS 2026 · 被引用 1 次
相关 Paper
- RecPipe: Co-designing Models and Hardware to Jointly Optimize Recommendation Quality and PerformanceUdit Gupta, Samuel Hsia, Jeff Zhang, Mark Wilkening 等MICRO 2021 · 被引用 31 次
- Pushing the Performance Envelope of DNN-based Recommendation Systems Inference on GPUsRishabh Jain, Vivek M. Bhasi, Adwait Jog, Anand Sivasubramaniam 等MICRO 2024 · 被引用 5 次
- Fleche: an efficient GPU embedding cache for personalized recommendationsMinhui Xie, Youyou Lu, Jiazhen Lin, Qing Wang 等EuroSys 2022 · 被引用 24 次
- AStitch: enabling a new multi-dimensional optimization space for memory-intensive ML training and inference on modern SIMT architecturesZhen Zheng, Xuanda Yang, Pengzhan Zhao, Guoping Long 等ASPLOS 2022 · 被引用 78 次
- Tensor Casting: Co-Designing Algorithm-Architecture for Personalized Recommendation TrainingYoungeun Kwon, Yunjae Lee, Minsoo RhuHPCA 2021 · 被引用 40 次
