BSL: Understanding and Improving Softmax Loss for Recommendation
Junkang Wu, Jiawei Chen, Jiancan Wu, Wentao Shi, Jizhi Zhang, Xiang Wang
摘要
Loss functions steer the optimization direction of recommendation models and are critical to model performance, but have received relatively little attention in recent recommendation research. Among various losses, we find Softmax loss (SL) stands out for not only achieving remarkable accuracy but also better robustness and fairness. Nevertheless, the current literature lacks a comprehensive explanation for the efficacy of SL. Toward addressing this research gap, we conduct theoretical analyses on SL and uncover three insights: 1) Optimizing SL is equivalent to performing Distributionally Robust Optimization (DRO) on the negative data, thereby learning against perturbations on the negative distribution and yielding robustness to noisy negatives. 2) Comparing with other loss functions, SL implicitly penalizes the prediction variance, resulting in a smaller gap between predicted values and and thus producing fairer results. Building on these insights, we further propose a novel loss function Bilateral SoftMax Loss (BSL) that extends the advantage of SL to both positive and negative sides. BSL augments SL by applying the same Log-Expectation-Exp structure to positive examples as is used for negatives, making the model robust to the noisy positives as well. Remarkably, BSL is simple and easy-to-implement - requiring just one additional line of code compared to SL. Experiments on four real-world datasets and three representative backbones demonstrate the effectiveness of our proposal. The code is available at https://github.com/junkangwu/BSL.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Distributionally Robust Graph-based Recommendation SystemBohao Wang, Jiawei Chen, Changdong Li, Sheng Zhou 等WWW 2024 · 被引用 42 次
- SIGformer: Sign-aware Graph Transformer for RecommendationSirui Chen, Jiawei Chen, Sheng Zhou, Bohao Wang 等SIGIR 2024 · 被引用 35 次
- PSL: Rethinking and Improving Softmax Loss from Pairwise Perspective for RecommendationWeiqin Yang, Jiawei Chen, Xin Xin, Sheng Zhou 等NeurIPS 2024 · 被引用 17 次
- Uncovering the Propensity Identification Problem in Debiased RecommendationsHonglei Zhang, Shuyi Wang, Haoxuan Li, Chunyuan Zheng 等ICDE 2024 · 被引用 12 次
- To Cool or not to Cool? Temperature Network Meets Large Foundation Models via DROZi-Hao Qiu, Siqi Guo, Mao Xu, Tuo Zhao 等ICML 2024 · 被引用 11 次
它引用的顶会 Paper24
- LightGCN: Simplifying and Powering Graph Convolution Network for RecommendationXiangnan He, Kuan Deng, Xiang Wang, Yan Li 等SIGIR 2020 · 被引用 4,448 次
- Self-supervised Graph Learning for RecommendationJiancan Wu, Xiang Wang, Fuli Feng, Xiangnan He 等SIGIR 2021 · 被引用 1,476 次
- Are Graph Augmentations Necessary?: Simple Graph Contrastive Learning for RecommendationJunliang Yu, Hongzhi Yin, Xin Xia, Tong Chen 等SIGIR 2022 · 被引用 658 次
- Revisiting Graph Based Collaborative Filtering: A Linear Residual Graph Convolutional Network ApproachLei Chen, Le Wu, Richang Hong, Kun Zhang 等AAAI 2020 · 被引用 634 次
- Disentangled Graph Collaborative FilteringXiang Wang, Hongye Jin, An Zhang, Xiangnan He 等SIGIR 2020 · 被引用 621 次
相关 Paper
- Advancing Loss Functions in Recommender Systems: A Comparative Study with a Rényi Divergence-Based SolutionShengjia Zhang, Jiawei Chen, Changdong Li, Sheng Zhou 等AAAI 2025 · 被引用 6 次
- Label Distributionally Robust Losses for Multi-class Classification: Consistency, Robustness and AdaptivityDixian Zhu, Yiming Ying, Tianbao YangICML 2023 · 被引用 15 次
- When AUC meets DRO: Optimizing Partial AUC for Deep Learning with Non-Convex Convergence GuaranteeDixian Zhu, Gang Li, Bokun Wang, Xiaodong Wu 等ICML 2022 · 被引用 42 次
- Distributionally Robust Finetuning BERT for Covariate Drift in Spoken Language UnderstandingSamuel Broscheit, Quynh Do, Judith GaspersACL 2022
- -Softmax: Approximating One-Hot Vectors for Mitigating Label NoiseJialiang Wang, Xiong Zhou, Deming Zhai, Junjun Jiang 等NeurIPS 2024 · 被引用 10 次
