Ekko: A Large-Scale Deep Learning Recommender System with Low-Latency Model Update
Chijun Sima, Yao Fu, Man-Kit Sit, Liyi Guo, Xuri Gong, Feng Lin, Junyu Wu, Yongsheng Li, Haidong Rong, Pierre-Louis Aublin, Luo Mai
摘要
Deep Learning Recommender Systems (DLRSs) need to update models at low latency, thus promptly serving new users and content. Existing DLRSs, however, fail to do so. They train/validate models offline and broadcast entire models to global inference clusters. They thus incur significant model update latency (e.g. dozens of minutes), which adversely affects Service-Level Objectives (SLOs).
This paper describes Ekko, a novel DLRS that enables low-latency model updates. Its design idea is to allow model updates to be immediately disseminated to all inference clusters, thus bypassing long-latency model checkpoint, validation and broadcast. To realise this idea, we first design an efficient peer-to-peer model update dissemination algorithm. This algorithm exploits the sparsity and temporal locality in updating DLRS models to improve the throughput and latency of updating models. Further, Ekko has a model update scheduler that can prioritise, over busy networks, the sending of model updates that can largely affect SLOs. Finally, Ekko has an inference model state manager which monitors the SLOs of inference models and rollbacks the models if SLOdetrimental biased updates are detected. Evaluation results show that Ekko is orders of magnitude faster than state-ofthe-art DLRS systems. Ekko has been deployed in production for more than one year, serves over a billion users daily and reduces the model update latency compared to state-of-the-art systems from dozens of minutes to 2.4 seconds.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- Robust Preference-Guided Denoising for Graph based Social RecommendationYuhan Quan, Jingtao Ding, Chen Gao, Lingling Yi 等WWW 2023 · 被引用 85 次
- Dynamically Expandable Graph Convolution for Streaming RecommendationBowei He, Xu He, Yingxue Zhang, Ruiming Tang 等WWW 2023 · 被引用 60 次
- AdaEmbed: Adaptive Embedding for Large-Scale Recommendation ModelsFan Lai, Wei Zhang, Rui Liu, William Tsai 等OSDI 2023 · 被引用 23 次
- GPU-Disaggregated Serving for Deep Learning Recommendation Models at ScaleLingyun Yang, Yongchen Wang, Yinghao Yu, Qizhen Weng 等NSDI 2025 · 被引用 22 次
- Read-ME: Refactorizing LLMs as Router-Decoupled Mixture of Experts with System Co-DesignRuisi Cai, Yeonju Ro, Geon-Woo Kim, Peihao Wang 等NeurIPS 2024 · 被引用 21 次
它引用的顶会 Paper12
- On the Convergence of FedAvg on Non-IID DataXiang Li, Kaixuan Huang, Wenhao Yang, Shusen Wang 等ICLR 2020 · 被引用 2,930 次
- Serving DNNs like Clockwork: Performance Predictability from the Bottom UpArpan Gujarati, Reza Karimi, Safya Alzayat, Wei Hao 等OSDI 2020 · 被引用 392 次
- Pollux: Co-adaptive Cluster Scheduling for Goodput-Optimized Deep LearningAurick Qiao, Sang Keun Choe, Suhas Jayaram Subramanya, Willie Neiswanger 等OSDI 2021 · 被引用 258 次
- Sundial: Fault-tolerant Clock Synchronization for DatacentersYuliang Li, Gautam Kumar, Hema Hariharan, Hassan M. G. Wassel 等OSDI 2020 · 被引用 66 次
- Agile and Accurate CTR Prediction Model Training for Massive-Scale Online Advertising SystemsZhiqiang Xu, Dong Li, Weijie Zhao, Xing Shen 等SIGMOD 2021 · 被引用 38 次
相关 Paper
- QuickUpdate: a Real-Time Personalization System for Large-Scale Recommendation ModelsKiran Kumar Matam, Hani Ramezani, Fan Wang, Zeliang Chen 等NSDI 2024 · 被引用 13 次
- Near-Zero-Overhead Freshness for Recommendation Systems via Inference-Side Model UpdatesWenjun Yu, Sitian Chen, Cheng Chen, Amelie Chi ZhouHPCA 2026 · 被引用 1 次
- Distributed Equivalent Substitution Training for Large-Scale Recommender SystemsHaidong Rong, Yangzihao Wang, Feihu Zhou, Junjie Zhai 等SIGIR 2020 · 被引用 9 次
- Accelerating Neural Recommendation Training with Embedding SchedulingChaoliang Zeng, Xudong Liao, Xiaodian Cheng, Han Tian 等NSDI 2024 · 被引用 16 次
- HypeReca: Distributed Heterogeneous In-Memory Embedding Database for Training Recommender ModelsJiaao He, Shengqi Chen, Kezhao Huang, Jidong ZhaiUSENIX ATC 2025 · 被引用 2 次
