Diversified Batch Selection for Training Acceleration
Feng Hong, Yueming Lyu, Jiangchao Yao, Ya Zhang, Ivor W. Tsang, Yanfeng Wang
Abstract
The remarkable success of modern machine learning models on large datasets often demands extensive training time and resource consumption. To save cost, a prevalent research line, known as online batch selection, explores selecting informative subsets during the training process. Although recent efforts achieve advancements by measuring the impact of each sample on generalization, their reliance on additional reference models inherently limits their practical applications, when there are no such ideal models available. On the other hand, the vanilla reference-model-free methods involve independently scoring and selecting data in a sample-wise manner, which sacrifices the diversity and induces the redundancy. To tackle this dilemma, we propose Diversified Batch Selection (DivBS), which is reference-model-free and can efficiently select diverse and representative samples. Specifically, we define a novel selection objective that measures the group-wise orthogonalized representativeness to combat the redundancy issue of previous sample-wise criteria, and provide a principled selection-efficient realization. Extensive experiments across various tasks demonstrate the significant superiority of DivBS in the performance-speedup trade-off. The code is publicly available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers14
- Revive Re-weighting in Imbalanced Learning by Density Ratio EstimationJiaan Luo, Feng Hong, Jiangchao Yao, Bo Han et al.NeurIPS 2024 · 16 citations
- Evolution-aware VAriance (EVA) Coreset Selection for Medical Image ClassificationYuxin Hong, Xiao Zhang, Xin Zhang, Joey Tianyi ZhouACM MM 2024 · 4 citations
- Partial Forward Blocking: A Novel Data Pruning Paradigm for Lossless Training AccelerationDongyue Wu, Zilin Guo, Jialong Zuo, Nong Sang et al.ICCV 2025 · 2 citations
- Reinforcement Learning-Guided Data Selection Via Redundancy AssessmentSuorong Yang, Peijia Li, Furao Shen, Jian ZhaoICCV 2025 · 1 citation
- Accelerating Benchmarking of Functional Connectivity Modeling via Structure-aware Core-set SelectionLing Zhan, Zhen Li, Junjie Huang, Tao JiaICLR 2026 · 1 citation
Builds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- Long-tail learning via logit adjustmentAditya Krishna Menon, Sadeep Jayasumana, Ankit Singh Rawat, Himanshu Jain et al.ICLR 2021 · 937 citations
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 806 citations
Related papers
- GREATS: Online Selection of High-Quality Data for LLM Training in Every IterationJiachen T. Wang, Tong Wu, Dawn Song, Prateek Mittal et al.NeurIPS 2024 · 91 citations
- Towards Accelerated Model Training via Bayesian Data SelectionZhijie Deng, Peng Cui, Jun ZhuNeurIPS 2023 · 15 citations
- Utility-Diversity Aware Online Batch Selection for LLM Supervised Fine-tuningHeming Zou, Yixiu Mao, Yun Qu, Qi Wang et al.ICML 2026 · 13 citations
- REDUCR: Robust Data Downsampling using Class Priority ReweightingWilliam Bankes, George Hughes, Ilija Bogunovic, Zi WangNeurIPS 2024 · 5 citations
- Leveraging Importance Weights in Subset SelectionGui Citovsky, Giulia DeSalvo, Sanjiv Kumar, Srikumar Ramalingam et al.ICLR 2023
