On the Role of Weight Decay in Collaborative Filtering: A Popularity Perspective
Donald Loveland, Mingxuan Ju, Tong Zhao, Neil Shah, Danai Koutra
Abstract
Collaborative filtering (CF) enables large-scale recommendation systems by encoding information from historical user-item interactions into dense ID-embedding tables. However, as embedding tables grow, closed-form solutions become impractical, often necessitating the use of mini-batch gradient descent for training. Despite extensive work on designing loss functions to train CF models, we argue that one core component of these pipelines is heavily overlooked: weight decay. Attaining high-performing models typically requires careful tuning of weight decay, regardless of loss, yet its necessity is not well understood. In this work, we question why weight decay is crucial in CF pipelines and how it impacts training. Through theoretical and empirical analysis, we surprisingly uncover that weight decay's primary function is to encode popularity information into the magnitudes of the embedding vectors. Moreover, we find that tuning weight decay acts as a coarse, non-linear, knob to influence preference towards popular or unpopular items. Based on these findings, we propose PRISM (Popularity-awaRe Initialization Strategy for embedding Magnitudes), a straightforward yet effective solution to simplify the training of high-performing CF models. PRISM pre-encodes the popularity information typically learned through weight decay, eliminating its necessity. Our experiments show that PRISM improves performance by up to 4.77% and reduces training times by 38.48%, compared to state-of-the-art training strategies. Additionally, we parameterize PRISM to modulate the initialization strength, offering a cost-effective and meaningful strategy to mitigate popularity bias.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 88c311b9-b9f6-4130-a70e-d30ff678cb8aCited by top-tier papers2
- AgentDR: Dynamic Recommendation with Implicit Item-Item Relations via LLM-based AgentsMingdai Yang, Nurendra Choudhary, Jiangshu Du, Edward W. Huang et al.WWW 2026
- Rethinking Popularity Bias in Collaborative Filtering via Analytical Vector DecompositionLingfeng Liu, Yixin Song, Dazhong Shen, Bing Yin et al.KDD 2026
Builds on10
- LightGCN: Simplifying and Powering Graph Convolution Network for RecommendationXiangnan He, Kuan Deng, Xiang Wang, Yan Li et al.SIGIR 2020 · 4,448 citations
- DCN V2: Improved Deep & Cross Network and Practical Lessons for Web-scale Learning to Rank SystemsRuoxi Wang, Rakesh Shivanna, Derek Zhiyuan Cheng, Sagar Jain et al.WWW 2021 · 793 citations
- Model-Agnostic Counterfactual Reasoning for Eliminating Popularity Bias in Recommender SystemTianxin Wei, Fuli Feng, Jiawei Chen, Ziwei Wu et al.KDD 2021 · 246 citations
- Towards Representation Alignment and Uniformity in Collaborative FilteringChenyang Wang, Yuanqing Yu, Weizhi Ma, Min Zhang et al.KDD 2022 · 179 citations
- AutoDebias: Learning to Debias for RecommendationJiawei Chen, Hande Dong, Yang Qiu, Xiangnan He et al.SIGIR 2021 · 167 citations
Related papers
- Understanding and Scaling Collaborative Filtering Optimization from the Perspective of Matrix RankDonald Loveland, Xinyi Wu, Tong Zhao, Danai Koutra et al.WWW 2025 · 9 citations
- Post-hoc Popularity Bias Correction in GNN-based Collaborative FilteringMd Aminul Islam, Elena Zheleva, Ren WangWWW 2026
- Mitigating the Popularity Bias of Graph Collaborative Filtering: A Dimensional Collapse PerspectiveYifei Zhang, Hao Zhu, Yankai Chen, Zixing Song et al.NeurIPS 2023 · 47 citations
- Causal Intervention for Leveraging Popularity Bias in RecommendationYang Zhang, Fuli Feng, Xiangnan He, Tianxin Wei et al.SIGIR 2021 · 431 citations
- Does Weighting Improve Matrix Factorization for Recommender Systems?Alex Ayoub, Samuel Robertson, Dawen Liang, Harald Steck et al.WWW 2025
