Near-Zero-Overhead Freshness for Recommendation Systems via Inference-Side Model Updates
Wenjun Yu, Sitian Chen, Cheng Chen, Amelie Chi Zhou
Abstract
Deep Learning Recommendation Models (DLRMs) underpin personalized services but face a critical freshnessaccuracy tradeoff due to massive parameter synchronization overheads. Production DLRMs deploy decoupled training/inference clusters, where synchronizing petabyte-scale embedding tables (EMTs) causes multi-minute staleness, degrading recommendation quality and revenue. We observe that (1) inference nodes exhibit sustained CPU underutilization (peak), and (2) EMT gradients possess intrinsic low-rank structure, enabling compact update representation. We present LiveUpdate, a system that eliminates inter-cluster synchronization by colocating Low-Rank Adaptation (LoRA) trainers within inference nodes. LiveUpdate addresses two core challenges: (1) dynamic rank adaptation via singular value monitoring to constrain memory overhead (of EMTs), and (2) NUMA-aware resource scheduling with hardware-enforced QoS to eliminate updateinference contention (P99 latency impact). Evaluations show LiveUpdate reduces update costs byversus delta-update baselines while achieving higher accuracy within 1-hour windows. By transforming idle inference resources into freshness engines, LiveUpdate delivers online model updates while outperforming state- of-the-art delta-update methods byin accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ce18d519-80bb-4028-af76-977dab9fe322Builds on16
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- FedPara: Low-rank Hadamard Product for Communication-Efficient Federated LearningNam Hyeon-Woo, Moon Ye-Bin, Tae-Hyun OhICLR 2022 · 179 citations
- Parameter-Efficient Fine-Tuning with Discrete Fourier TransformZiqi Gao, Qichao Wang, Aochuan Chen, Zijing Liu et al.ICML 2024 · 71 citations
- HET: Scaling out Huge Embedding Model Training via Cache-enabled Distributed FrameworkXupeng Miao, Hailin Zhang, Yining Shi, Xiaonan Nie et al.VLDB 2022 · 70 citations
- RecShard: statistical feature-based memory optimization for industry-scale neural recommendationGeet Sethi, Bilge Acun, Niket Agarwal, Christos Kozyrakis et al.ASPLOS 2022 · 65 citations
Related papers
- QuickUpdate: a Real-Time Personalization System for Large-Scale Recommendation ModelsKiran Kumar Matam, Hani Ramezani, Fan Wang, Zeliang Chen et al.NSDI 2024 · 13 citations
- Ekko: A Large-Scale Deep Learning Recommender System with Low-Latency Model UpdateChijun Sima, Yao Fu, Man-Kit Sit, Liyi Guo et al.OSDI 2022 · 31 citations
- Accelerating Neural Recommendation Training with Embedding SchedulingChaoliang Zeng, Xudong Liao, Xiaodian Cheng, Han Tian et al.NSDI 2024 · 16 citations
- Load and MLP-Aware Thread Orchestration for Recommendation Systems Inference on CPUsRishabh Jain, Teyuh Chou, Onur Kayiran, John Kalamatianos et al.ASPLOS 2025 · 4 citations
- Flora: Low-Rank Adapters Are Secretly Gradient CompressorsYongchang Hao, Yanshuai Cao, Lili MouICML 2024 · 113 citations
