Principled Synthetic Data Enables the First Scaling Laws for LLMs in Recommendation
Benyu Zhang, Qiang Zhang, Jianpeng Cheng, Hong-You Chen, Qifei wang, Wei Sun, Shen Li, Jia Li, Jiahao Wu, Xiangjun Fan, Hong Yan
Abstract
Large Language Models (LLMs) represent a promising frontier for recommender systems, yet their development has been impeded by the absence of predictable scaling laws, which are crucial for guiding research and optimizing resource allocation. We hypothesize that this may be attributed to the inherent noise, bias, and incompleteness of raw user interaction data in prior continual pre-training (CPT) efforts. This paper introduces a novel, layered framework for generating high-quality synthetic data that circumvents such issues by creating a curated, pedagogical curriculum for the LLM. We provide powerful, direct evidence for the utility of our curriculum by showing that standard sequential models trained on our principled synthetic data significantly outperform ( on recall@100 for SasRec) models trained on real data in downstream ranking tasks, demonstrating its superiority for learning generalizable user preference patterns. Building on this, we empirically demonstrate, for the first time, robust power-law scaling for an LLM that is continually pre-trained on our high-quality, recommendation-specific data. Our experiments reveal consistent and predictable perplexity reduction across multiple synthetic data modalities. These findings establish a foundational methodology for reliable scaling LLM capabilities in the recommendation domain, thereby shifting the research focus from mitigating data deficiencies to leveraging high-quality, structured information.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3c081ea1-0f82-4bab-9b11-ee87c5fe9884Builds on12
- Physics of Language Models: Part 3.1, Knowledge Storage and ExtractionZeyuan Allen-Zhu, Yuanzhi LiICML 2024 · 258 citations
- Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative RecommendationsJiaqi Zhai, Lucy Liao, Xing Liu, Yueming Wang et al.ICML 2024 · 200 citations
- OneRec-Think: In-Text Reasoning for Generative RecommendationZhanyu Liu, Shiyao Wang, Xingmei Wang, Rongzhou Zhang et al.ACL 2026 · 48 citations
- Rethinking Bias Mitigation: Fairer Architectures Make for Fairer Face RecognitionSamuel Dooley, Rhea Sanjay Sukthanker, John P. Dickerson, Colin White et al.NeurIPS 2023 · 41 citations
- BR-SNIS: Bias Reduced Self-Normalized Importance SamplingGabriel Cardoso, Sergey Samsonov, Achille Thin, Eric Moulines et al.NeurIPS 2022 · 21 citations
Related papers
- Perplexity-Aware Data Scaling Law: Perplexity Landscapes Predict Performance for Continual Pre-trainingLei Liu, Hao Zhu, Xiaoyan Yang, Yue Shen et al.ACL 2026
- Lost in Sequence: Do Large Language Models Understand Sequential Recommendation?Sein Kim, Hongseok Kang, Kibum Kim, Jiwan Kim et al.KDD 2025 · 3 citations
- CMR Scaling Law: Predicting Critical Mixture Ratios for Continual Pre-training of Language ModelsJiawei Gu, Zacc Yang, Chuanghao Ding, Rui Zhao et al.EMNLP 2024 · 2 citations
- LLM4RSR: Large Language Models as Data Correctors for Robust Sequential RecommendationYatong Sun, Xiaochun Yang, Zhu Sun, Yan Wang et al.AAAI 2025 · 2 citations
- Generative Archetype-Grounded Item Representations for Sequential RecommendationYifan Li, Jiahong Liu, Xinni Zhang, Hao Chen et al.WWW 2026
