Data Pruning via Moving-one-Sample-out
Haoru Tan, Sitong Wu, Fei Du, Yukang Chen, Zhibin Wang, Fan Wang, Xiaojuan Qi
Abstract
In this paper, we propose a novel data-pruning approach called moving-one-sampleout (MoSo), which aims to identify and remove the least informative samples from the training set. The core insight behind MoSo is to determine the importance of each sample by assessing its impact on the optimal empirical risk. This is achieved by measuring the extent to which the empirical risk changes when a particular sample is excluded from the training set. Instead of using the computationally expensive leaving-one-out-retraining procedure, we propose an efficient first-order approximator that only requires gradient information from different training stages. The key idea behind our approximation is that samples with gradients that are consistently aligned with the average gradient of the training set are more informative and should receive higher scores, which could be intuitively understood as follows: if the gradient from a specific sample is consistent with the average gradient vector, it implies that optimizing the network using the sample will yield a similar effect on all remaining samples. Experimental results demonstrate that MoSo effectively mitigates severe performance degradation at high pruning ratios and achieves satisfactory performance across various settings. Existing approaches can be broadly categorized into three major groups: pruning by importance criteria [11, 42, 28, 27, 26, 46] , coverage or diversity-driven methods [48, 29, 33, 15] , and optimizationbased methods [38, 8, 23, 24, 21, 44] . Among these, the first group of studies is the most effective and popular. These studies assume that hard samples are critical and informative core-set samples, and thus, they design difficulty-based metrics to assess sample importance. Such metrics include prediction entropy [11] , forgetting [28] or memorization [46] score, gradient norm [27], E2LN (variance of prediction) [27], self-supervised prototype distance [42], diverse ensembles [26], and others.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 806450c2-a89a-4607-851f-b03f78ae5f6cCited by top-tier papers36
- Data-efficient Fine-tuning for LLM-based RecommendationXinyu Lin, Wenjie Wang, Yongqi Li, Shuo Yang et al.SIGIR 2024 · 152 citations
- Mind the Boundary: Coreset Selection via Reconstructing the Decision BoundaryShuo Yang, Zhe Cao, Sheng Guo, Ruiheng Zhang et al.ICML 2024 · 25 citations
- On the Diversity and Realism of Distilled Dataset: An Efficient Dataset Distillation ParadigmPeng Sun, Bei Shi, Daiwei Yu, Tao LinCVPR 2024 · 17 citations
- How Far are AI-Generated Videos from Simulating the 3D Visual World: A Learned 3D Evaluation ApproachChirui Chang, Jiahui Liu, Zhengzhe Liu, Xiaoyang Lyu et al.ICCV 2025 · 15 citations
- Spanning Training Progress: Temporal Dual-Depth Scoring (TDDS) for Enhanced Dataset PruningXin Zhang, Jiawei Du, Yunsong Li, Weiying Xie et al.CVPR 2024 · 12 citations
Builds on19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu et al.NeurIPS 2022 · 2,727 citations
- Beyond neural scaling laws: beating power law scaling via data pruningBen Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli et al.NeurIPS 2022 · 720 citations
Related papers
- Batch Loss Score for Dynamic Data PruningQing Zhou, Bingxuan Zhao, Tao Yang, Hongyuan Zhang et al.CVPR 2026
- Dataset Pruning: Reducing Training Data by Examining Generalization InfluenceShuo Yang, Zeke Xie, Hanyu Peng, Min Xu et al.ICLR 2023 · 21 citations
- UNSEEN: Enhancing Dataset Pruning from a Generalization PerspectiveFurui Xu, Shaobo Wang, Jiajun Zhang, Chenghao Sun et al.AAAI 2026
- Selecting Samples on Graphs: A Unified Dataset Pruning Framework for Lossless Training AccelerationDongyue Wu, Zilin Guo, Xiaoyu Li, Jiajia Liu et al.ICML 2026
- Data Pruning by Information MaximizationHaoru Tan, Sitong Wu, Wei Huang, Shizhen Zhao et al.ICLR 2025
