Lune

HPCA2026顶会

Near-Zero-Overhead Freshness for Recommendation Systems via Inference-Side Model Updates

Wenjun Yu, Sitian Chen, Cheng Chen, Amelie Chi Zhou

2026年份
1被引次数

摘要

Deep Learning Recommendation Models (DLRMs) underpin personalized services but face a critical freshnessaccuracy tradeoff due to massive parameter synchronization overheads. Production DLRMs deploy decoupled training/inference clusters, where synchronizing petabyte-scale embedding tables (EMTs) causes multi-minute staleness, degrading recommendation quality and revenue. We observe that (1) inference nodes exhibit sustained CPU underutilization (peak≤20%\leq 20 \%), and (2) EMT gradients possess intrinsic low-rank structure, enabling compact update representation. We present LiveUpdate, a system that eliminates inter-cluster synchronization by colocating Low-Rank Adaptation (LoRA) trainers within inference nodes. LiveUpdate addresses two core challenges: (1) dynamic rank adaptation via singular value monitoring to constrain memory overhead (<2%<2 \%of EMTs), and (2) NUMA-aware resource scheduling with hardware-enforced QoS to eliminate updateinference contention (P99 latency impact<20 ms<20 ~\text{ms}). Evaluations show LiveUpdate reduces update costs by2×2 \timesversus delta-update baselines while achieving higher accuracy within 1-hour windows. By transforming idle inference resources into freshness engines, LiveUpdate delivers online model updates while outperforming state- of-the-art delta-update methods by0.04−0.24%\mathbf{0. 0 4 - 0. 2 4 \%}in accuracy.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper16

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖