Scale Efficient Training for Large Datasets
Qing Zhou, Junyu Gao, Qi Wang
Abstract
The rapid growth of dataset scales has been a key driver in advancing deep learning research. However, as dataset scale increases, the training process becomes increasingly inefficient due to the presence of low-value samples, including excessive redundant samples, overly challenging samples, and inefficient easy samples that contribute little to model improvement. To address this challenge, we propose Scale Efficient Training (SeTa) for large datasets, a dynamic sample pruning approach that losslessly reduces training time. To remove low-value samples, SeTa first performs random pruning to eliminate redundant samples, then clusters the remaining samples according to their learning difficulty measured by loss. Building upon this clustering, a sliding window strategy is employed to progressively remove both overly challenging and inefficient easy clusters following an easy-to-hard curriculum. We conduct extensive experiments on large-scale synthetic datasets, including ToCa, SS1M, and ST+MJ, each containing over 3 million samples. SeTa reduces training costs by up to 50% while maintaining or improving performance, with minimal degradation even at 70% cost reduction. Furthermore, experiments on various scale real datasets across various backbones (CNNs, Transformers, and Mambas) and diverse tasks (instruction tuning, multi-view stereo, geo-localization, composed image retrieval, referring image segmentation) demonstrate the powerful effectiveness and universality of our approach. Code is available at https://github.com/mrazhou/SeTa .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8e8d78db-3efa-4754-b4cc-b021542409fdCited by top-tier papers4
- Beyond Prompt Degradation: Prototype-guided Dual-pool Prompting for Incremental Object DetectionYaoteng Zhang, Qing Zhou, Junyu Gao, Qi WangCVPR 2026 · 2 citations
- Learnability-Guided Diffusion for Dataset DistillationJeffrey A. Chan-Santiago, Mubarak ShahCVPR 2026 · 2 citations
- Batch Loss Score for Dynamic Data PruningQing Zhou, Bingxuan Zhao, Tao Yang, Hongyuan Zhang et al.CVPR 2026
- Inconsistency Biases in Dynamic Data PruningQing Zhou, Tao Yang, Bingxuan Zhao, Hongyuan Zhang et al.ICLR 2026
Builds on18
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- EfficientNetV2: Smaller Models and Faster TrainingMingxing Tan, Quoc V. LeICML 2021 · 4,239 citations
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang et al.ICML 2024 · 1,725 citations
Related papers
- Repeated Random Sampling for Minimizing the Time-to-Accuracy of LearningPatrik Okanovic, Roger Waleffe, Vasilis Mageirakos, Konstantinos E. Nikolakakis et al.ICLR 2024 · 29 citations
- Primitive3D: 3D Object Dataset Synthesis from Randomly Assembled PrimitivesXinke Li, Henghui Ding, Zekun Tong, Yuwei Wu et al.CVPR 2022 · 7 citations
- Improving the Scaling Laws of Synthetic Data with Deliberate PracticeReyhane Askari Hemmat, Mohammad Pezeshki, Elvis Dohmatob, Florian Bordes et al.ICML 2025
- Exploring Learning Complexity for Efficient Downstream Dataset PruningWenyu Jiang, Zhenlong Liu, Zejian Xie, Songxin Zhang et al.ICLR 2025
- Effective pruning of web-scale datasets based on complexity of concept clustersAmro Abbas, Evgenia Rusak, Kushal Tirumala, Wieland Brendel et al.ICLR 2024 · 30 citations
