When Newer is Not Better: Does Deep Learning Really Benefit Recommendation From Implicit Feedback?
Yushun Dong, Jundong Li, Tobias Schnabel
Abstract
In recent years, neural models have been repeatedly touted to exhibit state-of-the-art performance in recommendation. Nevertheless, multiple recent studies have revealed that the reported stateof-the-art results of many neural recommendation models cannot be reliably replicated. A primary reason is that existing evaluations are performed under various inconsistent protocols. Correspondingly, these replicability issues make it difficult to understand how much benefit we can actually gain from these neural models. It then becomes clear that a fair and comprehensive performance comparison between traditional and neural models is needed.
Motivated by these issues, we perform a large-scale, systematic study to compare recent neural recommendation models against traditional ones in top-𝑛 recommendation from implicit data. We propose a set of evaluation strategies for measuring memorization performance, generalization performance, and subgroup-specific performance of recommendation models. We conduct extensive experiments with 13 popular recommendation models (including two neural models and 11 traditional ones as baselines) on nine commonly used datasets. Our experiments demonstrate that even with extensive hyper-parameter searches, neural models do not dominate traditional models in all aspects, e.g., they fare worse in terms of average HitRate. We further find that there are areas where neural models seem to outperform non-neural models, for example, in recommendation diversity and robustness between different subgroups of users and items. Our work illuminates the relative advantages and disadvantages of neural models in recommendation and is therefore an important step towards building better recommender systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9ffe17bf-579e-43e2-874a-3f1790c7e2d5Cited by top-tier papers2
- Why is Normalization Necessary for Linear Recommenders?Seongmin Park, Mincheol Yoon, Hye-young Kim, Jongwuk LeeSIGIR 2025 · 1 citation
- Quantifying User Coherence: A Unified Framework for Analyzing Recommender Systems Across DomainsMichaël Soumm, Alexandre Fournier-Montgieux, Adrian Popescu, Bertrand DelezoideWWW 2026
Builds on5
- LightGCN: Simplifying and Powering Graph Convolution Network for RecommendationXiangnan He, Kuan Deng, Xiang Wang, Yan Li et al.SIGIR 2020 · 4,448 citations
- Where to Go Next: Modeling Long- and Short-Term User Preferences for Point-of-Interest RecommendationKe Sun, Tieyun Qian, Tong Chen, Yile Liang et al.AAAI 2020 · 412 citations
- User-oriented Fairness in RecommendationYunqi Li, Hanxiong Chen, Zuohui Fu, Yingqiang Ge et al.WWW 2021 · 293 citations
- On Sampling Top-K Recommendation EvaluationDong Li, Ruoming Jin, Jing Gao, Zhi LiuKDD 2020 · 43 citations
- Towards a Better Understanding of Linear Models for RecommendationRuoming Jin, Dong Li, Jing Gao, Zhi Liu et al.KDD 2021 · 21 citations
Related papers
- Are Neural Rankers still Outperformed by Gradient Boosted Decision Trees?Zhen Qin, Le Yan, Honglei Zhuang, Yi Tay et al.ICLR 2021 · 41 citations
- On the Generalizability and Predictability of Recommender SystemsDuncan C. McElfresh, Sujay Khandagale, Jonathan Valverde, John Dickerson et al.NeurIPS 2022 · 17 citations
- SetRank: A Setwise Bayesian Approach for Collaborative Ranking from Implicit FeedbackChao Wang, Hengshu Zhu, Chen Zhu, Chuan Qin et al.AAAI 2020 · 53 citations
- Efficient Heterogeneous Collaborative Filtering without Negative Sampling for RecommendationChong Chen, Min Zhang, Yongfeng Zhang, Weizhi Ma et al.AAAI 2020 · 185 citations
- Does Weighting Improve Matrix Factorization for Recommender Systems?Alex Ayoub, Samuel Robertson, Dawen Liang, Harald Steck et al.WWW 2025
