Load and MLP-Aware Thread Orchestration for Recommendation Systems Inference on CPUs
Rishabh Jain, Teyuh Chou, Onur Kayiran, John Kalamatianos, Gabriel H. Loh, Mahmut T. Kandemir, Chita R. Das
Abstract
Recommendation models can enhance consumer experiences and are one of the most frequently used machine learning models in data centers. The deep learning recommendation model (DLRM) is one such key workload. While DLRMs are often trained using GPUs, CPUs can be a cost-effective solution for inference. Therefore, optimizing DLRM inference for CPUs is an important research problem with significant business value. In this work, we identify several shortcomings of existing DLRM parallelization techniques, which can include load imbalance across CPU chiplets, suboptimal core allocation for embedding tables, and inefficient utilization of memory- level parallelism (MLP) resources. We propose a novel thread scheduler, called ''Balance,'' that addresses those shortcomings by (1) minimizing core allocation per embedding table to maximize core utilization, (2) using MLP-aware task scheduling based on the characteristics of the embedding tables to better utilize memory bandwidth, and (3) combining work stealing and table reordering mechanisms to reduce load imbalance across CPU chiplets. We evaluate Balance on real hardware with production DLRM traces and demonstrate up to a 1.67× higher speedup over prior state-of-the-art DLRM parallelization techniques with 96 cores. Further, Balance consistently achieves 1.22× higher performance over a range of batch sizes.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get f3e8684f-ef0b-42d2-89e0-c0bedb4068a7Cited by top-tier papers1
Ask how each one uses itRelated papers
- UpDLRM: Accelerating Personalized Recommendation using Real-World PIM ArchitectureSitian Chen, Haobin Tan, Amelie Chi Zhou, Yusen Li et al.DAC 2024 · 9 citations
- Optimizing CPU Performance for Recommendation Systems At-ScaleRishabh Jain, Scott Cheng, Vishwas Kalagi, Vrushabh Sanghavi et al.ISCA 2023 · 25 citations
- RecNMP: Accelerating Personalized Recommendation with Near-Memory ProcessingLiu Ke, Udit Gupta, Benjamin Youngjae Cho, David Brooks et al.ISCA 2020 · 235 citations
- RecFlow: Unlocking GPU Efficiency for DLRM Inference via Fine-Grained Parallelism and Incremental BatchingSiheng Pan, Shaolong Li, Minwei Zhang, Shuxi Guo et al.INFOCOM 2026
- Accelerating Neural Recommendation Training with Embedding SchedulingChaoliang Zeng, Xudong Liao, Xiaodian Cheng, Han Tian et al.NSDI 2024 · 16 citations
