Beyond Dense Connectivity: Explicit Sparsity for Scalable Recommendation
Yantao Yu, Sen Qiao, Lei Shen, Bing Wang, Xiaoyi Zeng
摘要
Recent progress in scaling large models has motivated recommender systems to increase model depth and capacity to better leverage massive behavioral data. However, recommendation inputs are high-dimensional and extremely sparse, and simply scaling dense backbones (e.g., deep MLPs) often yields diminishing returns or even performance degradation. Our analysis of industrial CTR models reveals a phenomenon of implicit connection sparsity: most learned connection weights tend towards zero, while only a small fraction remain prominent. This indicates a structural mismatch between dense connectivity and sparse recommendation data; by compelling the model to process vast low-utility connections instead of valid signals, the dense architecture itself becomes the primary bottleneck to effective pattern modeling. We propose SSR (Explicit Sparsity for Scalable Recommendation), a framework that incorporates sparsity explicitly into the architecture. SSR employs a multi-view ''filter-then-fuse'' mechanism, decomposing inputs into parallel views for dimension-level sparse filtering followed by dense fusion. Specifically, we realize the sparsity via two strategies: a Static Random Filter that achieves efficient structural sparsity via fixed dimension subsets, and Iterative Competitive Sparse (ICS), a differentiable dynamic mechanism that employs bio-inspired competition to adaptively retain high-response dimensions. Experiments on three public datasets and a billion-scale industrial dataset from AliExpress (a global e-commerce platform) show that SSR outperforms state-of-the-art baselines under similar budgets. Crucially, SSR exhibits superior scalability, delivering continuous performance gains where dense models saturate. The code is available at https://github.com/Atticus666/SSRNet.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer 等NeurIPS 2021 · 被引用 3,862 次
- DCN V2: Improved Deep & Cross Network and Practical Lessons for Web-scale Learning to Rank SystemsRuoxi Wang, Rakesh Shivanna, Derek Zhiyuan Cheng, Sagar Jain 等WWW 2021 · 被引用 793 次
- Adaptive Factorization Network: Learning Adaptive-Order Feature InteractionsWeiyu Cheng, Yanyan Shen, Linpeng HuangAAAI 2020 · 被引用 202 次
- Wukong: Towards a Scaling Law for Large-Scale RecommendationBuyun Zhang, Liang Luo, Yuxin Chen, Jade Nie 等ICML 2024 · 被引用 108 次
- Are wider nets better given the same number of parameters?Anna Golubeva, Guy Gur-Ari, Behnam NeyshaburICLR 2021 · 被引用 48 次
相关 Paper
- Distributed Equivalent Substitution Training for Large-Scale Recommender SystemsHaidong Rong, Yangzihao Wang, Feihu Zhou, Junjie Zhai 等SIGIR 2020 · 被引用 9 次
- Can Recommender Systems Teach Themselves? A Recursive Self-Improving Framework with Fidelity ControlLuankang Zhang, Hao Wang, Zhongzhou Liu, MINGJIA YIN 等ICML 2026 · 被引用 5 次
- AutoDim: Field-aware Embedding Dimension Searchin Recommender SystemsXiangyu Zhao, Haochen Liu, Hui Liu, Jiliang Tang 等WWW 2021 · 被引用 69 次
- NASRec: Weight Sharing Neural Architecture Search for Recommender SystemsTunhou Zhang, Dehua Cheng, Yuchen He, Zhengxing Chen 等WWW 2023 · 被引用 20 次
- Scaling Sequential Recommendation Models with TransformersPablo Zivic, Hernán Ceferino Vázquez, Jorge SánchezSIGIR 2024 · 被引用 25 次
