Efficient Content-based Recommendation Model Training via Noise-aware Coreset Selection
Hung Vinh Tran, Tong Chen, Hechuan Wen, Quoc Viet Hung Nguyen, Bin Cui, Hongzhi Yin
Abstract
Content-based recommendation systems (CRSs) utilize content features to predict user-item interactions, serving as essential tools for helping users navigate information-rich web services. However, ensuring the effectiveness of CRSs requires large-scale and even continuous model training to accommodate diverse user preferences, resulting in significant computational costs and resource demands. A promising approach to this challenge is coreset selection, which identifies a small but representative subset of data samples that preserves model quality while reducing training overhead. Yet, the selected coreset is vulnerable to the pervasive noise in user-item interactions, particularly when it is minimally sized. To this end, we propose Noise-aware Coreset Selection (NaCS), a specialized framework for CRSs. NaCS constructs coresets through submodular optimization based on training gradients, while simultaneously correcting noisy labels using a progressively trained model. Meanwhile, we refine the selected coreset by filtering out low-confidence samples through uncertainty quantification, thereby avoid training with unreliable interactions. Through extensive experiments, we show that NaCS produces higher-quality coresets for CRSs while achieving better efficiency than existing coreset selection techniques. Notably, NaCS recovers 93-95% of full-dataset training performance using merely 1% of the training data. The source code is available at https://github.com/chenxing1999/nacs . CCS Concepts • Information systems → Recommender systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b033a9ac-395a-4e7d-a93f-f9a7199cd4aeBuilds on31
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 806 citations
- DCN V2: Improved Deep & Cross Network and Practical Lessons for Web-scale Learning to Rank SystemsRuoxi Wang, Rakesh Shivanna, Derek Zhiyuan Cheng, Sagar Jain et al.WWW 2021 · 793 citations
- Coresets for Data-efficient Training of Machine Learning ModelsBaharan Mirzasoleiman, Jeff A. Bilmes, Jure LeskovecICML 2020 · 494 citations
- Selection via Proxy: Efficient Data Selection for Deep LearningCody Coleman, Christopher Yeh, Stephen Mussmann, Baharan Mirzasoleiman et al.ICLR 2020 · 462 citations
- Coresets via Bilevel Optimization for Continual Learning and StreamingZalán Borsos, Mojmir Mutny, Andreas KrauseNeurIPS 2020 · 320 citations
Related papers
- Efficient Representativeness-Aware Coreset SelectionZihao Cheng, Binrui Wu, Zhiwei Li, Yuesen Liao et al.NeurIPS 2025 · 1 citation
- Efficient Adversarial Contrastive Learning via Robustness-Aware Coreset SelectionXilie Xu, Jingfeng Zhang, Feng Liu, Masashi Sugiyama et al.NeurIPS 2023 · 26 citations
- Efficient Coreset Selection with Cluster-based MethodsChengliang Chai, Jiayi Wang, Nan Tang, Ye Yuan et al.KDD 2023 · 19 citations
- CRSSC: Salvage Reusable Samples from Noisy Data for Robust LearningZeren Sun, Xian-Sheng Hua, Yazhou Yao, Xiu-Shen Wei et al.ACM MM 2020 · 57 citations
- Is Bin Generation Indispensable? A Bin-Generation-Free Dataset Quantization via Semantic PerspectiveMaijie Deng, Yuhua Li, Yixiong Zou, Yao Wu et al.CVPR 2026
