Optimizing CPU Performance for Recommendation Systems At-Scale
Rishabh Jain, Scott Cheng, Vishwas Kalagi, Vrushabh Sanghavi, Samvit Kaul, Meena Arunachalam, Kiwan Maeng, Adwait Jog, Anand Sivasubramaniam, Mahmut Taylan Kandemir, Chita R. Das
Abstract
Deep Learning Recommendation Models (DLRMs) are very popular in personalized recommendation systems and are a major contributor to the data-center AI cycles. Due to the high computational and memory bandwidth needs of DLRMs, specifically the embedding stage in DLRM inferences, both CPUs and GPUs are used for hosting such workloads. This is primarily because of the heavy irregular memory accesses in the embedding stage of computation that leads to significant stalls in the CPU pipeline. As the model and parameter sizes keep increasing with newer recommendation models, the computational dominance of the embedding stage also grows, thereby, bringing into question the suitability of CPUs for inference. In this paper, we first quantify the cause of irregular accesses and their impact on caches and observe that off-chip memory access is the main contributor to high latency. Therefore, we exploit two well-known techniques: (1) Software prefetching, to hide the memory access latency suffered by the demand loads and (2) Overlapping computation and memory accesses, to reduce CPU stalls via hyperthreading to minimize the overall execution time. We evaluate our work on a single-core and 24-core configuration with the latest recommendation models and recently released production traces. Our integrated techniques speed up the inference by up to 1.59x, and on average by 1.4x.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 436fde99-d4ae-49af-b77c-3ae93370d1f1Cited by top-tier papers9
- Managing Memory Tiers with CXL in Virtualized EnvironmentsYuhong Zhong, Daniel S. Berger, Carl A. Waldspurger, Ryan Wee et al.OSDI 2024 · 77 citations
- GPU-Disaggregated Serving for Deep Learning Recommendation Models at ScaleLingyun Yang, Yongchen Wang, Yinghao Yu, Qizhen Weng et al.NSDI 2025 · 22 citations
- UpDLRM: Accelerating Personalized Recommendation using Real-World PIM ArchitectureSitian Chen, Haobin Tan, Amelie Chi Zhou, Yusen Li et al.DAC 2024 · 9 citations
- PIFS-Rec: Process-In-Fabric-Switch for Large-Scale Recommendation System InferencesPingyi Huo, Anusha Devulapally, Hasan Al Maruf, Minseo Park et al.MICRO 2024 · 6 citations
- Machine Learning-Guided Memory Optimization for DLRM Inference on Tiered MemoryJie Ren, Bin Ma, Shuangyan Yang, Benjamin Francis et al.HPCA 2025 · 6 citations
Related papers
- Pushing the Performance Envelope of DNN-based Recommendation Systems Inference on GPUsRishabh Jain, Vivek M. Bhasi, Adwait Jog, Anand Sivasubramaniam et al.MICRO 2024 · 5 citations
- Load and MLP-Aware Thread Orchestration for Recommendation Systems Inference on CPUsRishabh Jain, Teyuh Chou, Onur Kayiran, John Kalamatianos et al.ASPLOS 2025 · 4 citations
- RecNMP: Accelerating Personalized Recommendation with Near-Memory ProcessingLiu Ke, Udit Gupta, Benjamin Youngjae Cho, David Brooks et al.ISCA 2020 · 235 citations
- RecFlow: Unlocking GPU Efficiency for DLRM Inference via Fine-Grained Parallelism and Incremental BatchingSiheng Pan, Shaolong Li, Minwei Zhang, Shuxi Guo et al.INFOCOM 2026
- Training personalized recommendation systems from (GPU) scratch: look forward not backwardsYoungeun Kwon, Minsoo RhuISCA 2022 · 24 citations
