BSL: Understanding and Improving Softmax Loss for Recommendation
Junkang Wu, Jiawei Chen, Jiancan Wu, Wentao Shi, Jizhi Zhang, Xiang Wang
Abstract
Loss functions steer the optimization direction of recommendation models and are critical to model performance, but have received relatively little attention in recent recommendation research. Among various losses, we find Softmax loss (SL) stands out for not only achieving remarkable accuracy but also better robustness and fairness. Nevertheless, the current literature lacks a comprehensive explanation for the efficacy of SL. Toward addressing this research gap, we conduct theoretical analyses on SL and uncover three insights: 1) Optimizing SL is equivalent to performing Distributionally Robust Optimization (DRO) on the negative data, thereby learning against perturbations on the negative distribution and yielding robustness to noisy negatives. 2) Comparing with other loss functions, SL implicitly penalizes the prediction variance, resulting in a smaller gap between predicted values and and thus producing fairer results. Building on these insights, we further propose a novel loss function Bilateral SoftMax Loss (BSL) that extends the advantage of SL to both positive and negative sides. BSL augments SL by applying the same Log-Expectation-Exp structure to positive examples as is used for negatives, making the model robust to the noisy positives as well. Remarkably, BSL is simple and easy-to-implement - requiring just one additional line of code compared to SL. Experiments on four real-world datasets and three representative backbones demonstrate the effectiveness of our proposal. The code is available at https://github.com/junkangwu/BSL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 68427e57-da05-49e0-a245-3e33d46b6c00Cited by top-tier papers11
- Distributionally Robust Graph-based Recommendation SystemBohao Wang, Jiawei Chen, Changdong Li, Sheng Zhou et al.WWW 2024 · 42 citations
- SIGformer: Sign-aware Graph Transformer for RecommendationSirui Chen, Jiawei Chen, Sheng Zhou, Bohao Wang et al.SIGIR 2024 · 35 citations
- PSL: Rethinking and Improving Softmax Loss from Pairwise Perspective for RecommendationWeiqin Yang, Jiawei Chen, Xin Xin, Sheng Zhou et al.NeurIPS 2024 · 17 citations
- Uncovering the Propensity Identification Problem in Debiased RecommendationsHonglei Zhang, Shuyi Wang, Haoxuan Li, Chunyuan Zheng et al.ICDE 2024 · 12 citations
- To Cool or not to Cool? Temperature Network Meets Large Foundation Models via DROZi-Hao Qiu, Siqi Guo, Mao Xu, Tuo Zhao et al.ICML 2024 · 11 citations
Builds on24
- LightGCN: Simplifying and Powering Graph Convolution Network for RecommendationXiangnan He, Kuan Deng, Xiang Wang, Yan Li et al.SIGIR 2020 · 4,448 citations
- Self-supervised Graph Learning for RecommendationJiancan Wu, Xiang Wang, Fuli Feng, Xiangnan He et al.SIGIR 2021 · 1,476 citations
- Are Graph Augmentations Necessary?: Simple Graph Contrastive Learning for RecommendationJunliang Yu, Hongzhi Yin, Xin Xia, Tong Chen et al.SIGIR 2022 · 658 citations
- Revisiting Graph Based Collaborative Filtering: A Linear Residual Graph Convolutional Network ApproachLei Chen, Le Wu, Richang Hong, Kun Zhang et al.AAAI 2020 · 634 citations
- Disentangled Graph Collaborative FilteringXiang Wang, Hongye Jin, An Zhang, Xiangnan He et al.SIGIR 2020 · 621 citations
Related papers
- Advancing Loss Functions in Recommender Systems: A Comparative Study with a Rényi Divergence-Based SolutionShengjia Zhang, Jiawei Chen, Changdong Li, Sheng Zhou et al.AAAI 2025 · 6 citations
- Label Distributionally Robust Losses for Multi-class Classification: Consistency, Robustness and AdaptivityDixian Zhu, Yiming Ying, Tianbao YangICML 2023 · 15 citations
- When AUC meets DRO: Optimizing Partial AUC for Deep Learning with Non-Convex Convergence GuaranteeDixian Zhu, Gang Li, Bokun Wang, Xiaodong Wu et al.ICML 2022 · 42 citations
- Distributionally Robust Finetuning BERT for Covariate Drift in Spoken Language UnderstandingSamuel Broscheit, Quynh Do, Judith GaspersACL 2022
- -Softmax: Approximating One-Hot Vectors for Mitigating Label NoiseJialiang Wang, Xiong Zhou, Deming Zhai, Junjun Jiang et al.NeurIPS 2024 · 10 citations
