Towards Accelerated Model Training via Bayesian Data Selection
Zhijie Deng, Peng Cui, Jun Zhu
Abstract
Mislabeled, duplicated, or biased data in real-world scenarios can lead to prolonged training and even hinder model convergence. Traditional solutions prioritizing easy or hard samples lack the flexibility to handle such a variety simultaneously. Recent work has proposed a more reasonable data selection principle by examining the data's impact on the model's generalization loss. However, its practical adoption relies on less principled approximations and additional holdout data. This work solves these problems by leveraging a lightweight Bayesian treatment and incorporating off-the-shelf zero-shot predictors built on large-scale pre-trained models. The resulting algorithm is efficient and easy to implement. We perform extensive empirical studies on challenging benchmarks with considerable data noise and imbalance in the online batch selection scenario, and observe superior training efficiency over competitive baselines. Notably, on the challenging Web-Vision benchmark, our method can achieve similar predictive performance with significantly fewer training iterations than leading data selection methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 74dacc01-65e3-4e91-989a-8e168b3733d5Cited by top-tier papers9
- GREATS: Online Selection of High-Quality Data for LLM Training in Every IterationJiachen T. Wang, Tong Wu, Dawn Song, Prateek Mittal et al.NeurIPS 2024 · 91 citations
- Adaptive Defense against Harmful Fine-Tuning for Large Language Models via Bayesian Data SchedulerZixuan Hu, Li Shen, Zhenyi Wang, Yongxian Wei et al.NeurIPS 2025 · 16 citations
- Diversified Batch Selection for Training AccelerationFeng Hong, Yueming Lyu, Jiangchao Yao, Ya Zhang et al.ICML 2024 · 16 citations
- Utility-Diversity Aware Online Batch Selection for LLM Supervised Fine-tuningHeming Zou, Yixiu Mao, Yun Qu, Qi Wang et al.ICML 2026 · 13 citations
- Partial Forward Blocking: A Novel Data Pruning Paradigm for Lossless Training AccelerationDongyue Wu, Zilin Guo, Jialong Zuo, Nong Sang et al.ICCV 2025 · 2 citations
Builds on14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 806 citations
Related papers
- Navigating Towards Fairness with Data SelectionYixuan Zhang, Zhidong Li, Yang Wang, Fang Chen et al.AAAI 2025 · 1 citation
- REDUCR: Robust Data Downsampling using Class Priority ReweightingWilliam Bankes, George Hughes, Ilija Bogunovic, Zi WangNeurIPS 2024 · 5 citations
- Prioritized Training on Points that are Learnable, Worth Learning, and not yet LearntSören Mindermann, Jan Markus Brauner, Muhammed Razzak, Mrinank Sharma et al.ICML 2022 · 237 citations
- GLISTER: Generalization based Data Subset Selection for Efficient and Robust LearningKrishnaTeja Killamsetty, Durga Sivasubramanian, Ganesh Ramakrishnan, Rishabh K. IyerAAAI 2021 · 300 citations
- When Dynamic Data Selection Meets Data Augmentation: Achieving Enhanced Training AccelerationSuorong Yang, Peng Ye, Furao Shen, Dongzhan ZhouICML 2025
