CAFE: Towards Compact, Adaptive, and Fast Embedding for Large-scale Recommendation Models
Hailin Zhang, Zirui Liu, Boxuan Chen, Yikai Zhao, Tong Zhao, Tong Yang, Bin Cui
摘要
Recently, the growing memory demands of embedding tables in Deep Learning Recommendation Models (DLRMs) pose great challenges for model training and deployment. Existing embedding compression solutions cannot simultaneously meet three key design requirements: memory efficiency, low latency, and adaptability to dynamic data distribution. This paper presents CAFE, a Compact, Adaptive, and Fast Embedding compression framework that addresses the above requirements. The design philosophy of CAFE is to dynamically allocate more memory resources to important features (called hot features), and allocate less memory to unimportant ones. In CAFE, we propose a fast and lightweight sketch data structure, named HotSketch, to capture feature importance and report hot features in real time. For each reported hot feature, we assign it a unique embedding. For the non-hot features, we allow multiple features to share one embedding by using hash embedding technique. Guided by our design philosophy, we further propose a multi-level hash embedding framework to optimize the embedding tables of non-hot features. We theoretically analyze the accuracy of HotSketch, and analyze the model convergence against deviation. Extensive experiments show that CAFE significantly outperforms existing embedding compression methods, yielding 3.92% and 3.68% superior testing AUC on Criteo Kaggle dataset and CriteoTB dataset at a compression ratio of 10000×. The source codes of CAFE are available at GitHub [75] .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language ModelsXin Cheng, Wangding Zeng, Damai Dai, Qinyu Chen 等ACL 2026 · 被引用 57 次
- Surge Phenomenon in Optimal Learning Rate and Batch Size ScalingShuaipeng Li, Penghao Zhao, Hailin Zhang, Xingwu Sun 等NeurIPS 2024 · 被引用 33 次
- PQCache: Product Quantization-based KVCache for Long Context LLM InferenceHailin Zhang, Xiaodong Ji, Yilin Chen, Fangcheng Fu 等SIGMOD 2025 · 被引用 13 次
- OPER: Optimality-Guided Embedding Table Parallelization for Large-scale Recommendation ModelZheng Wang, Yuke Wang, Boyuan Feng, Guyue Huang 等USENIX ATC 2024 · 被引用 7 次
- PIFS-Rec: Process-In-Fabric-Switch for Large-Scale Recommendation System InferencesPingyi Huo, Anusha Devulapally, Hasan Al Maruf, Minseo Park 等MICRO 2024 · 被引用 6 次
它引用的顶会 Paper24
- DeepRecSys: A System for Optimizing End-To-End At-Scale Neural Recommendation InferenceUdit Gupta, Samuel Hsia, Vikram Saraph, Xiaodong Wang 等ISCA 2020 · 被引用 149 次
- CocoSketch: high-performance sketch-based measurement over arbitrary partial key queryYinda Zhang, Zaoxing Liu, Ruixin Wang, Tong Yang 等SIGCOMM 2021 · 被引用 146 次
- QueryFormer: A Tree Transformer Model for Query Plan RepresentationYue Zhao, Gao Cong, Jiachen Shi, Chunyan MiaoVLDB 2022 · 被引用 117 次
- WavingSketch: An Unbiased and Generic Sketch for Finding Top-k Items in Data StreamsJizhou Li, Zikun Li, Yifei Xu, Shiqi Jiang 等KDD 2020 · 被引用 96 次
- Compositional Embeddings Using Complementary Partitions for Memory-Efficient Recommendation SystemsHao-Jun Michael Shi, Dheevatsa Mudigere, Maxim Naumov, Jiyan YangKDD 2020 · 被引用 88 次
相关 Paper
- The trade-offs of model size in large recommendation models : 100GB to 10MB Criteo-tb DLRM modelAditya Desai, Anshumali ShrivastavaNeurIPS 2022 · 被引用 17 次
- Balanced Co-Clustering of Users and Items for Embedding Table Compression in Recommender SystemsRunhao Jiang, Renchi Yang, Donghao WuSIGIR 2026
- AdaEmbed: Adaptive Embedding for Large-Scale Recommendation ModelsFan Lai, Wei Zhang, Rui Liu, William Tsai 等OSDI 2023 · 被引用 23 次
- Hybrid Embedding Framework for Memory-Efficient Recommendation SystemsSeung Jin Yang, Hyuk-Jae Lee, Chae-Eun RheeDAC 2025
- Clustering the Sketch: Dynamic Compression for Embedding TablesHenry Ling-Hei Tsang, Thomas D. AhleNeurIPS 2023 · 被引用 5 次
