Wukong: Towards a Scaling Law for Large-Scale Recommendation
Buyun Zhang, Liang Luo, Yuxin Chen, Jade Nie, Xi Liu, Shen Li, Yanli Zhao, Yuchen Hao, Yantao Yao, Ellie Dingqiao Wen, Jongsoo Park, Maxim Naumov, Wenlin Chen
Abstract
Scaling laws play an instrumental role in the sustainable improvement in model quality. Unfortunately, recommendation models to date do not exhibit such laws similar to those observed in the domain of large language models, due to the inefficiencies of their upscaling mechanisms. This limitation poses significant challenges in adapting these models to increasingly more complex real-world datasets. In this paper, we propose an effective network architecture based purely on stacked factorization machines, and a synergistic upscaling strategy, collectively dubbed Wukong, to establish a scaling law in the domain of recommendation. Wukong's unique design makes it possible to capture diverse, any-order of interactions simply through taller and wider layers. We conducted extensive evaluations on six public datasets, and our results demonstrate that Wukong consistently outperforms state-of-the-art models quality-wise. Further, we assessed Wukong's scalability on an internal, large-scale dataset. The results show that Wukong retains its superiority in quality over state-of-the-art models, while holding the scaling law across two orders of magnitude in model complexity, extending beyond 100 GFLOP/example, where prior arts fall short.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 153d70f5-9a9a-4079-ace1-9f8981709aa1Cited by top-tier papers15
- P-Law: Predicting Quantitative Scaling Law with Entropy Guidance in Large Recommendation ModelsTingjia Shen, Hao Wang, Chuhan Wu, Jin Yao Chin et al.NeurIPS 2025 · 9 citations
- Phantora: Maximizing Code Reuse in Simulation-based Machine Learning System Performance EstimationJianxing Qin, Jingrong Chen, Xinhao Kong, Yongji Wu et al.NSDI 2026 · 5 citations
- Field Matters: A Lightweight LLM-enhanced Method for CTR PredictionYu Cui, Feng Liu, Jiawei Chen, Xingyu Lou et al.WWW 2026 · 5 citations
- HyFormer: Revisiting the Roles of Sequence Modeling and Feature Interaction in CTR PredictionYunwen Huang, Shiyong Hong, Xijun Xiao, Jinqiu Jin et al.SIGIR 2026 · 4 citations
- Primus: Unified Training System for Large-Scale Deep Learning Recommendation ModelsJixi Shan, Xiuqi Huang, Yang Guo, Hongyue Mao et al.USENIX ATC 2025 · 4 citations
Builds on5
- DCN V2: Improved Deep & Cross Network and Practical Lessons for Web-scale Learning to Rank SystemsRuoxi Wang, Rakesh Shivanna, Derek Zhiyuan Cheng, Sagar Jain et al.WWW 2021 · 793 citations
- FinalMLP: An Enhanced Two-Stream MLP Model for CTR PredictionKelong Mao, Jieming Zhu, Liangcai Su, Guohao Cai et al.AAAI 2023 · 142 citations
- Scaling Law for Recommendation Models: Towards General-Purpose User RepresentationsKyuyong Shin, Hanock Kwak, Su Young Kim, Max Nihlén Ramström et al.AAAI 2023 · 57 citations
- On the Embedding Collapse when Scaling up Recommendation ModelsXingzhuo Guo, Junwei Pan, Ximei Wang, Baixu Chen et al.ICML 2024 · 55 citations
- Learning to Embed Categorical Features without Embedding Tables for RecommendationWang-Cheng Kang, Derek Zhiyuan Cheng, Tiansheng Yao, Xinyang Yi et al.KDD 2021 · 46 citations
Related papers
- Principled Synthetic Data Enables the First Scaling Laws for LLMs in RecommendationBenyu Zhang, Qiang Zhang, Jianpeng Cheng, Hong-You Chen et al.ICML 2026
- Breaking the Bottleneck: User-Specific Optimization and Real-Time Inference Integration for Sequential RecommendationWenjia Xie, Hao Wang, Minghao Fang, Ruize Yu et al.KDD 2025
- StackRec: Efficient Training of Very Deep Sequential Recommender Models by Iterative StackingJiachun Wang, Fajie Yuan, Jian Chen, Qingyao Wu et al.SIGIR 2021 · 25 citations
- Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative RecommendationsJiaqi Zhai, Lucy Liao, Xing Liu, Yueming Wang et al.ICML 2024 · 200 citations
- Expand More, Shrink Less: Shaping Effective-Rank Dynamics for Dense Scaling in RecommendationGuoming Li, Shangyu Zhang, Junwei Pan, Wentao Ning et al.KDD 2026 · 2 citations
