Hercules: Heterogeneity-Aware Inference Serving for At-Scale Personalized Recommendation
Liu Ke, Udit Gupta, Mark Hempstead, Carole-Jean Wu, Hsien-Hsin S. Lee, Xuan Zhang
Abstract
Personalized recommendation is an important class of deep-learning applications that powers a large collection of internet services and consumes a considerable amount of datacenter resources. As the scale of production-grade recommendation systems continues to grow, optimizing their serving performance and efficiency in a heterogeneous datacenter is important and can translate into infrastructure capacity saving. In this paper, we propose Hercules, an optimized framework for personalized recommendation inference serving that targets diverse industry-representative models and cloud-scale heterogeneous systems. Hercules performs a two-stage optimization procedure — offline profiling and online serving. The first stage searches the large under-explored task scheduling space with a gradient-based search algorithm achieving up to 9.0× latency-bounded throughput improvement on individual servers; it also identifies the optimal heterogeneous server architecture for each recommendation workload. The second stage performs heterogeneity-aware cluster provisioning to optimize resource mapping and allocation in response to fluctuating diurnal loads. The proposed cluster scheduler in Hercules achieves 47.7% cluster capacity saving and reduces the provisioned power by 23.7% over a state-of-the-art greedy scheduler.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- KRISP: Enabling Kernel-wise RIght-sizing for Spatial Partitioned GPU Inference ServersMarcus Chow, Ali Jahanshahi, Daniel WongHPCA 2023 · 25 citations
- GPU-Disaggregated Serving for Deep Learning Recommendation Models at ScaleLingyun Yang, Yongchen Wang, Yinghao Yu, Qizhen Weng et al.NSDI 2025 · 22 citations
- MP-Rec: Hardware-Software Co-design to Enable Multi-path RecommendationSamuel Hsia, Udit Gupta, Bilge Acun, Newsha Ardalani et al.ASPLOS 2023 · 10 citations
- PreSto: An In-Storage Data Preprocessing System for Training Recommendation ModelsYunjae Lee, Hyeseong Kim, Minsoo RhuISCA 2024 · 8 citations
- PIFS-Rec: Process-In-Fabric-Switch for Large-Scale Recommendation System InferencesPingyi Huo, Anusha Devulapally, Hasan Al Maruf, Minseo Park et al.MICRO 2024 · 6 citations
Builds on5
- RecNMP: Accelerating Personalized Recommendation with Near-Memory ProcessingLiu Ke, Udit Gupta, Benjamin Youngjae Cho, David Brooks et al.ISCA 2020 · 235 citations
- DeepRecSys: A System for Optimizing End-To-End At-Scale Neural Recommendation InferenceUdit Gupta, Samuel Hsia, Vikram Saraph, Xiaodong Wang et al.ISCA 2020 · 149 citations
- FAFNIR: Accelerating Sparse Gathering by Using Efficient Near-Memory Intelligent ReductionBahar Asgari, Ramyad Hadidi, Jiashen Cao, Da Eun Shim et al.HPCA 2021 · 87 citations
- Lazy Batching: An SLA-aware Batching System for Cloud Machine Learning InferenceYujeong Choi, Yunseong Kim, Minsoo RhuHPCA 2021 · 65 citations
- Tensor Casting: Co-Designing Algorithm-Architecture for Personalized Recommendation TrainingYoungeun Kwon, Yunjae Lee, Minsoo RhuHPCA 2021 · 40 citations
Related papers
- RecPipe: Co-designing Models and Hardware to Jointly Optimize Recommendation Quality and PerformanceUdit Gupta, Samuel Hsia, Jeff Zhang, Mark Wilkening et al.MICRO 2021 · 31 citations
- Power-aware Deep Learning Model Serving with μ-ServeHaoran Qiu, Weichao Mao, Archit Patke, Shengkun Cui et al.USENIX ATC 2024 · 82 citations
- ElasticRec: A Microservice-based Model Serving Architecture Enabling Elastic Resource Scaling for Recommendation ModelsYujeong Choi, Jiin Kim, Minsoo RhuISCA 2024 · 2 citations
- Serving Heterogeneous Machine Learning Models on Multi-GPU Servers with Spatio-Temporal SharingSeungbeom Choi, Sunho Lee, Yeonjae Kim, Jongse Park et al.USENIX ATC 2022 · 200 citations
- HexGen-3: A Fully Disaggregated LLM Serving Framework with Fine-Grained Heterogeneous Resource AutoscalingYouhe Jiang, Wenshuang Li, You Peng, Jintao Zhang et al.ICML 2026
