UMEC: Unified model and embedding compression for efficient recommendation systems
Jiayi Shen, Haotao Wang, Shupeng Gui, Jianchao Tan, Zhangyang Wang, Ji Liu
Abstract
The recommendation system (RS) plays an important role in the content recommendation and retrieval scenarios. The core part of the system is the ranking neural network, which is usually a bottleneck of whole system performance during online inference. Hammering an efficient neural network-based recommendation system involves entangled challenges of compressing both the network parameters and the feature embedding inputs. We propose a unified model and embedding compression (UMEC) framework to jointly learn input feature selection and neural network compression together, which is formulated as a resource-constrained optimization problem and solved using the alternating direction method of multipliers (ADMM) algorithm. Experimental results on public benchmarks show that our UMEC framework notably outperforms other non-integrated baseline methods. The codes can be found at https://github.com/VITA-Group/UMEC .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ff634f0a-52e7-418e-acb3-60a32af7740fCited by top-tier papers6
- Unified Visual Transformer CompressionShixing Yu, Tianlong Chen, Jiayi Shen, Huan Yuan et al.ICLR 2022 · 118 citations
- GDP: Stabilized Neural Network Pruning via Gates with Differentiable PolarizationYi Guo, Huan Yuan, Jianchao Tan, Zhangyang Wang et al.ICCV 2021 · 52 citations
- Experimental Analysis of Large-scale Learnable Vector Storage CompressionHailin Zhang, Penghao Zhao, Xupeng Miao, Yingxia Shao et al.VLDB 2024 · 20 citations
- Resource Constrained Model Compression via Minimax Optimization for Spiking Neural NetworksJue Chen, Huan Yuan, Jianchao Tan, Bin Chen et al.ACM MM 2023 · 5 citations
- IDEAL: Query-Efficient Data-Free Learning from Black-Box ModelsJie Zhang, Chen Chen, Lingjuan LyuICLR 2023 · 5 citations
Builds on10
- Similarity-Preserving Knowledge DistillationFrederick Tung, Greg MoriICCV 2019 · 1,214 citations
- MetaPruning: Meta Learning for Automatic Neural Network Channel PruningZechun Liu, Haoyuan Mu, Xiangyu Zhang, Zichao Guo et al.ICCV 2019 · 633 citations
- The Lottery Ticket Hypothesis for Pre-trained BERT NetworksTianlong Chen, Jonathan Frankle, Shiyu Chang, Sijia Liu et al.NeurIPS 2020 · 428 citations
- A Unified Lottery Ticket Hypothesis for Graph Neural NetworksTianlong Chen, Yongduo Sui, Xuxi Chen, Aston Zhang et al.ICML 2021 · 208 citations
- AutoGAN-Distiller: Searching to Compress Generative Adversarial NetworksYonggan Fu, Wuyang Chen, Haotao Wang, Haoran Li et al.ICML 2020 · 91 citations
Related papers
- xLightFM: Extremely Memory-Efficient Factorization MachineGangwei Jiang, Hao Wang, Jin Chen, Haoyu Wang et al.SIGIR 2021 · 25 citations
- A Generic Network Compression Framework for Sequential Recommender SystemsYang Sun, Fajie Yuan, Min Yang, Guoao Wei et al.SIGIR 2020 · 52 citations
- CAFE: Towards Compact, Adaptive, and Fast Embedding for Large-scale Recommendation ModelsHailin Zhang, Zirui Liu, Boxuan Chen, Yikai Zhao et al.SIGMOD 2024 · 15 citations
- Towards Efficient Tensor Decomposition-Based DNN Model Compression With Optimization FrameworkMiao Yin, Yang Sui, Siyu Liao, Bo YuanCVPR 2021
- Unified Embedding: Battle-Tested Feature Representations for Web-Scale ML SystemsBenjamin Coleman, Wang-Cheng Kang, Matthew Fahrbach, Ruoxi Wang et al.NeurIPS 2023 · 30 citations
