Wukong: Towards a Scaling Law for Large-Scale Recommendation
Buyun Zhang, Liang Luo, Yuxin Chen, Jade Nie, Xi Liu, Shen Li, Yanli Zhao, Yuchen Hao, Yantao Yao, Ellie Dingqiao Wen, Jongsoo Park, Maxim Naumov, Wenlin Chen
摘要
Scaling laws play an instrumental role in the sustainable improvement in model quality. Unfortunately, recommendation models to date do not exhibit such laws similar to those observed in the domain of large language models, due to the inefficiencies of their upscaling mechanisms. This limitation poses significant challenges in adapting these models to increasingly more complex real-world datasets. In this paper, we propose an effective network architecture based purely on stacked factorization machines, and a synergistic upscaling strategy, collectively dubbed Wukong, to establish a scaling law in the domain of recommendation. Wukong's unique design makes it possible to capture diverse, any-order of interactions simply through taller and wider layers. We conducted extensive evaluations on six public datasets, and our results demonstrate that Wukong consistently outperforms state-of-the-art models quality-wise. Further, we assessed Wukong's scalability on an internal, large-scale dataset. The results show that Wukong retains its superiority in quality over state-of-the-art models, while holding the scaling law across two orders of magnitude in model complexity, extending beyond 100 GFLOP/example, where prior arts fall short.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- P-Law: Predicting Quantitative Scaling Law with Entropy Guidance in Large Recommendation ModelsTingjia Shen, Hao Wang, Chuhan Wu, Jin Yao Chin 等NeurIPS 2025 · 被引用 9 次
- Phantora: Maximizing Code Reuse in Simulation-based Machine Learning System Performance EstimationJianxing Qin, Jingrong Chen, Xinhao Kong, Yongji Wu 等NSDI 2026 · 被引用 5 次
- Field Matters: A Lightweight LLM-enhanced Method for CTR PredictionYu Cui, Feng Liu, Jiawei Chen, Xingyu Lou 等WWW 2026 · 被引用 5 次
- HyFormer: Revisiting the Roles of Sequence Modeling and Feature Interaction in CTR PredictionYunwen Huang, Shiyong Hong, Xijun Xiao, Jinqiu Jin 等SIGIR 2026 · 被引用 4 次
- Primus: Unified Training System for Large-Scale Deep Learning Recommendation ModelsJixi Shan, Xiuqi Huang, Yang Guo, Hongyue Mao 等USENIX ATC 2025 · 被引用 4 次
它引用的顶会 Paper5
- DCN V2: Improved Deep & Cross Network and Practical Lessons for Web-scale Learning to Rank SystemsRuoxi Wang, Rakesh Shivanna, Derek Zhiyuan Cheng, Sagar Jain 等WWW 2021 · 被引用 793 次
- FinalMLP: An Enhanced Two-Stream MLP Model for CTR PredictionKelong Mao, Jieming Zhu, Liangcai Su, Guohao Cai 等AAAI 2023 · 被引用 142 次
- Scaling Law for Recommendation Models: Towards General-Purpose User RepresentationsKyuyong Shin, Hanock Kwak, Su Young Kim, Max Nihlén Ramström 等AAAI 2023 · 被引用 57 次
- On the Embedding Collapse when Scaling up Recommendation ModelsXingzhuo Guo, Junwei Pan, Ximei Wang, Baixu Chen 等ICML 2024 · 被引用 55 次
- Learning to Embed Categorical Features without Embedding Tables for RecommendationWang-Cheng Kang, Derek Zhiyuan Cheng, Tiansheng Yao, Xinyang Yi 等KDD 2021 · 被引用 46 次
相关 Paper
- Principled Synthetic Data Enables the First Scaling Laws for LLMs in RecommendationBenyu Zhang, Qiang Zhang, Jianpeng Cheng, Hong-You Chen 等ICML 2026
- Breaking the Bottleneck: User-Specific Optimization and Real-Time Inference Integration for Sequential RecommendationWenjia Xie, Hao Wang, Minghao Fang, Ruize Yu 等KDD 2025
- StackRec: Efficient Training of Very Deep Sequential Recommender Models by Iterative StackingJiachun Wang, Fajie Yuan, Jian Chen, Qingyao Wu 等SIGIR 2021 · 被引用 25 次
- Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative RecommendationsJiaqi Zhai, Lucy Liao, Xing Liu, Yueming Wang 等ICML 2024 · 被引用 200 次
- Expand More, Shrink Less: Shaping Effective-Rank Dynamics for Dense Scaling in RecommendationGuoming Li, Shangyu Zhang, Junwei Pan, Wentao Ning 等KDD 2026 · 被引用 2 次
