On the Embedding Collapse when Scaling up Recommendation Models
Xingzhuo Guo, Junwei Pan, Ximei Wang, Baixu Chen, Jie Jiang, Mingsheng Long
摘要
Recent advances in foundation models have led to a promising trend of developing large recommendation models to leverage vast amounts of available data. Still, mainstream models remain embarrassingly small in size and naïve enlarging does not lead to sufficient performance gain, suggesting a deficiency in the model scalability. In this paper, we identify the embedding collapse phenomenon as the inhibition of scalability, wherein the embedding matrix tends to occupy a low-dimensional subspace. Through empirical and theoretical analysis, we demonstrate a two-sided effect of feature interaction specific to recommendation models. On the one hand, interacting with collapsed embeddings restricts embedding learning and exacerbates the collapse issue. On the other hand, interaction is crucial in mitigating the fitting of spurious features as a scalability guarantee. Based on our analysis, we propose a simple yet effective multi-embedding design incorporating embedding-set-specific interaction modules to learn embedding sets with large diversity and thus reduce collapse. Extensive experiments demonstrate that this proposed design provides consistent scalability and effective collapse mitigation for various recommendation models. Code is available at this repository: https://github.com/thuml/Multi-Embedding.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Wukong: Towards a Scaling Law for Large-Scale RecommendationBuyun Zhang, Liang Luo, Yuxin Chen, Jade Nie 等ICML 2024 · 被引用 108 次
- Listwise Preference Diffusion Optimization for User Behavior Trajectories PredictionHongtao Huang, Chengkai Huang, Junda Wu, Tong Yu 等NeurIPS 2025 · 被引用 16 次
- Order-agnostic Identifier for Large Language Model-based Generative RecommendationXinyu Lin, Haihan Shi, Wenjie Wang, Fuli Feng 等SIGIR 2025 · 被引用 15 次
- Pre-train, Align, and Disentangle: Empowering Sequential Recommendation with Large Language ModelsYuhao Wang, Junwei Pan, Pengyue Jia, Wanyu Wang 等SIGIR 2025 · 被引用 8 次
- Adaptive Regularization for Large-Scale Sparse Feature Embedding ModelsMang Li, Wei LyuICLR 2026 · 被引用 4 次
它引用的顶会 Paper15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Fine-Tuning can Distort Pretrained Features and Underperform Out-of-DistributionAnanya Kumar, Aditi Raghunathan, Robbie Matthew Jones, Tengyu Ma 等ICLR 2022 · 被引用 911 次
- DCN V2: Improved Deep & Cross Network and Practical Lessons for Web-scale Learning to Rank SystemsRuoxi Wang, Rakesh Shivanna, Derek Zhiyuan Cheng, Sagar Jain 等WWW 2021 · 被引用 793 次
- Understanding Dimensional Collapse in Contrastive Self-supervised LearningLi Jing, Pascal Vincent, Yann LeCun, Yuandong TianICLR 2022 · 被引用 467 次
相关 Paper
- Curse of "Low" Dimensionality in Recommender SystemsNaoto Ohsaka, Riku TogashiSIGIR 2023 · 被引用 8 次
- Expand More, Shrink Less: Shaping Effective-Rank Dynamics for Dense Scaling in RecommendationGuoming Li, Shangyu Zhang, Junwei Pan, Wentao Ning 等KDD 2026 · 被引用 2 次
- Learnable Embedding sizes for Recommender SystemsSiyi Liu, Chen Gao, Yihong Chen, Depeng Jin 等ICLR 2021 · 被引用 97 次
- The Pitfall of Scaling Up: Uncovering and Mitigating Popularity Bias Amplification in Scaling Transformer-based RecommendersWeiqin Yang, Yue Pan, Chongming Gao, Sheng Zhou 等KDD 2026
- INMO: A Model-Agnostic and Scalable Module for Inductive Collaborative FilteringYunfan Wu, Qi Cao, Huawei Shen, Shuchang Tao 等SIGIR 2022 · 被引用 22 次
