UMEC: Unified model and embedding compression for efficient recommendation systems
Jiayi Shen, Haotao Wang, Shupeng Gui, Jianchao Tan, Zhangyang Wang, Ji Liu
摘要
The recommendation system (RS) plays an important role in the content recommendation and retrieval scenarios. The core part of the system is the ranking neural network, which is usually a bottleneck of whole system performance during online inference. Hammering an efficient neural network-based recommendation system involves entangled challenges of compressing both the network parameters and the feature embedding inputs. We propose a unified model and embedding compression (UMEC) framework to jointly learn input feature selection and neural network compression together, which is formulated as a resource-constrained optimization problem and solved using the alternating direction method of multipliers (ADMM) algorithm. Experimental results on public benchmarks show that our UMEC framework notably outperforms other non-integrated baseline methods. The codes can be found at https://github.com/VITA-Group/UMEC .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Unified Visual Transformer CompressionShixing Yu, Tianlong Chen, Jiayi Shen, Huan Yuan 等ICLR 2022 · 被引用 118 次
- GDP: Stabilized Neural Network Pruning via Gates with Differentiable PolarizationYi Guo, Huan Yuan, Jianchao Tan, Zhangyang Wang 等ICCV 2021 · 被引用 52 次
- Experimental Analysis of Large-scale Learnable Vector Storage CompressionHailin Zhang, Penghao Zhao, Xupeng Miao, Yingxia Shao 等VLDB 2024 · 被引用 20 次
- Resource Constrained Model Compression via Minimax Optimization for Spiking Neural NetworksJue Chen, Huan Yuan, Jianchao Tan, Bin Chen 等ACM MM 2023 · 被引用 5 次
- IDEAL: Query-Efficient Data-Free Learning from Black-Box ModelsJie Zhang, Chen Chen, Lingjuan LyuICLR 2023 · 被引用 5 次
它引用的顶会 Paper10
- Similarity-Preserving Knowledge DistillationFrederick Tung, Greg MoriICCV 2019 · 被引用 1,214 次
- MetaPruning: Meta Learning for Automatic Neural Network Channel PruningZechun Liu, Haoyuan Mu, Xiangyu Zhang, Zichao Guo 等ICCV 2019 · 被引用 633 次
- The Lottery Ticket Hypothesis for Pre-trained BERT NetworksTianlong Chen, Jonathan Frankle, Shiyu Chang, Sijia Liu 等NeurIPS 2020 · 被引用 428 次
- A Unified Lottery Ticket Hypothesis for Graph Neural NetworksTianlong Chen, Yongduo Sui, Xuxi Chen, Aston Zhang 等ICML 2021 · 被引用 208 次
- AutoGAN-Distiller: Searching to Compress Generative Adversarial NetworksYonggan Fu, Wuyang Chen, Haotao Wang, Haoran Li 等ICML 2020 · 被引用 91 次
相关 Paper
- xLightFM: Extremely Memory-Efficient Factorization MachineGangwei Jiang, Hao Wang, Jin Chen, Haoyu Wang 等SIGIR 2021 · 被引用 25 次
- A Generic Network Compression Framework for Sequential Recommender SystemsYang Sun, Fajie Yuan, Min Yang, Guoao Wei 等SIGIR 2020 · 被引用 52 次
- CAFE: Towards Compact, Adaptive, and Fast Embedding for Large-scale Recommendation ModelsHailin Zhang, Zirui Liu, Boxuan Chen, Yikai Zhao 等SIGMOD 2024 · 被引用 15 次
- Towards Efficient Tensor Decomposition-Based DNN Model Compression With Optimization FrameworkMiao Yin, Yang Sui, Siyu Liao, Bo YuanCVPR 2021
- Unified Embedding: Battle-Tested Feature Representations for Web-Scale ML SystemsBenjamin Coleman, Wang-Cheng Kang, Matthew Fahrbach, Ruoxi Wang 等NeurIPS 2023 · 被引用 30 次
