AdaEmbed: Adaptive Embedding for Large-Scale Recommendation Models
Fan Lai, Wei Zhang, Rui Liu, William Tsai, Xiaohan Wei, Yuxi Hu, Sabin Devkota, Jianyu Huang, Jongsoo Park, Xing Liu, Zeliang Chen, Ellie Wen
Abstract
Deep learning recommendation models (DLRMs) are using increasingly larger embedding tables to represent categorical sparse features such as video genres. Each embedding row of the table represents the trainable weight vector for a specific instance of that feature. While increasing the number of embedding rows typically improves model accuracy by considering more feature instances, it can lead to larger deployment costs and slower model execution.
Unlike existing efforts that primarily focus on optimizing DLRMs for the given embedding, we present a complementary system, AdaEmbed, to reduce the size of embeddings needed for the same DLRM accuracy via in-training embedding pruning. Our key insight is that the access patterns and weights of different embeddings are heterogeneous across embedding rows, and dynamically change over the training process, implying varying embedding importance with respect to model accuracy. However, identifying important embeddings and then enforcing pruning for modern DLRMs with up to billions of embeddings (terabytes) is challenging. Given the total embedding size, AdaEmbed considers embeddings with higher runtime access frequencies and larger training gradients to be more important, and it dynamically prunes less important embeddings at scale to automatically determine per-feature embeddings. Our evaluations in industrial settings show that AdaEmbed saves 35-60% embedding size needed in deployment and improves model execution speed by 11-34%, while achieving noticeable accuracy gains.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5be149f8-debf-4506-b786-756e0dc590bfCited by top-tier papers7
- GPU-Disaggregated Serving for Deep Learning Recommendation Models at ScaleLingyun Yang, Yongchen Wang, Yinghao Yu, Qizhen Weng et al.NSDI 2025 · 22 citations
- CAFE: Towards Compact, Adaptive, and Fast Embedding for Large-scale Recommendation ModelsHailin Zhang, Zirui Liu, Boxuan Chen, Yikai Zhao et al.SIGMOD 2024 · 15 citations
- Pushing the Performance Envelope of DNN-based Recommendation Systems Inference on GPUsRishabh Jain, Vivek M. Bhasi, Adwait Jog, Anand Sivasubramaniam et al.MICRO 2024 · 5 citations
- On-device Content-based Recommendation with Single-shot Embedding Pruning: A Cooperative Game PerspectiveHung Vinh Tran, Tong Chen, Guanhua Ye, Quoc Viet Hung Nguyen et al.WWW 2025 · 4 citations
- RecFlex: Enabling Feature Heterogeneity-Aware Optimization for Deep Recommendation Models with Flexible SchedulesZaifeng Pan, Zhen Zheng, Feng Zhang, Bing Xie et al.SC 2024 · 2 citations
Builds on19
- Movement Pruning: Adaptive Sparsity by Fine-TuningVictor Sanh, Thomas Wolf, Alexander M. RushNeurIPS 2020 · 656 citations
- A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU ClustersYimin Jiang, Yibo Zhu, Chang Lan, Bairen Yi et al.OSDI 2020 · 390 citations
- PatDNN: Achieving Real-Time DNN Execution on Mobile Devices with Pattern-based Weight PruningWei Niu, Xiaolong Ma, Sheng Lin, Shihao Wang et al.ASPLOS 2020 · 214 citations
- PipeSwitch: Fast Pipelined Context Switching for Deep Learning ApplicationsZhihao Bai, Zhen Zhang, Yibo Zhu, Xin JinOSDI 2020 · 152 citations
- The Generalization-Stability Tradeoff In Neural Network PruningBrian R. Bartoldson, Ari S. Morcos, Adrian Barbu, Gordon ErlebacherNeurIPS 2020 · 97 citations
Related papers
- Learnable Embedding sizes for Recommender SystemsSiyi Liu, Chen Gao, Yihong Chen, Depeng Jin et al.ICLR 2021 · 97 citations
- The trade-offs of model size in large recommendation models : 100GB to 10MB Criteo-tb DLRM modelAditya Desai, Anshumali ShrivastavaNeurIPS 2022 · 17 citations
- EVStore: Storage and Caching Capabilities for Scaling Embedding Tables in Deep Recommendation SystemsDaniar Heri Kurniawan, Ruipu Wang, Kahfi S. Zulkifli, Fandi A. Wiranata et al.ASPLOS 2023 · 16 citations
- Hybrid Embedding Framework for Memory-Efficient Recommendation SystemsSeung Jin Yang, Hyuk-Jae Lee, Chae-Eun RheeDAC 2025
- FusedRec: Fused Embedding Communication for Distributed Recommendation Training on GPUsXuanteng Huang, Fan Li, Riyang Hu, Jianchang Zhang et al.AAAI 2026 · 1 citation
