Adaptive Regularization for Large-Scale Sparse Feature Embedding Models
Mang Li, Wei Lyu
Abstract
The one-epoch overfitting problem has drawn widespread attention, especially in CTR and CVR estimation models in search, advertising, and recommendation domains. These models which rely heavily on large-scale sparse categorical features, often suffer a significant decline in performance when trained for multiple epochs. Although recent studies have proposed heuristic solutions, the fundamental cause of this phenomenon remains unclear. In this work, we present a theoretical explanation grounded in Rademacher complexity, supported by empirical experiments, to explain why overfitting occurs in models with large-scale sparse categorical features. Based on this analysis, we propose a regularization method that constrains the norm budget of embedding layers adaptively. Our approach not only prevents the severe performance degradation observed during multi-epoch training, but also improves model performance within a single epoch. This method has already been deployed in online production systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on5
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- On the Embedding Collapse when Scaling up Recommendation ModelsXingzhuo Guo, Junwei Pan, Ximei Wang, Baixu Chen et al.ICML 2024 · 55 citations
- Kraken: memory-efficient continual learning for large-scale real-time recommendationsMinhui Xie, Kai Ren, Youyou Lu, Guangxu Yang et al.SC 2020 · 38 citations
- Scaling Transformers for Discriminative Recommendation via Generative PretrainingChunqi Wang, Bingchao Wu, Zheng Chen, Lei Shen et al.KDD 2025 · 1 citation
Related papers
- CowClip: Reducing CTR Prediction Model Training Time from 12 Hours to 10 Minutes on 1 GPUZangwei Zheng, Pengtai Xu, Xuan Zou, Da Tang et al.AAAI 2023 · 9 citations
- Generalizable Multi-Pass Training of Ads Recommendation Models with Foundation Model GuidanceYunzhe Qi, Qinghai Zhou, Boyang Liu, Can Cui et al.KDD 2026
- Unified Embedding: Battle-Tested Feature Representations for Web-Scale ML SystemsBenjamin Coleman, Wang-Cheng Kang, Matthew Fahrbach, Ruoxi Wang et al.NeurIPS 2023 · 30 citations
- FM2: Field-matrixed Factorization Machines for Recommender SystemsYang Sun, Junwei Pan, Alex Zhang, Aaron FloresWWW 2021 · 98 citations
- AdaEmbed: Adaptive Embedding for Large-Scale Recommendation ModelsFan Lai, Wei Zhang, Rui Liu, William Tsai et al.OSDI 2023 · 23 citations
