Data Pruning via Moving-one-Sample-out
Haoru Tan, Sitong Wu, Fei Du, Yukang Chen, Zhibin Wang, Fan Wang, Xiaojuan Qi
摘要
In this paper, we propose a novel data-pruning approach called moving-one-sampleout (MoSo), which aims to identify and remove the least informative samples from the training set. The core insight behind MoSo is to determine the importance of each sample by assessing its impact on the optimal empirical risk. This is achieved by measuring the extent to which the empirical risk changes when a particular sample is excluded from the training set. Instead of using the computationally expensive leaving-one-out-retraining procedure, we propose an efficient first-order approximator that only requires gradient information from different training stages. The key idea behind our approximation is that samples with gradients that are consistently aligned with the average gradient of the training set are more informative and should receive higher scores, which could be intuitively understood as follows: if the gradient from a specific sample is consistent with the average gradient vector, it implies that optimizing the network using the sample will yield a similar effect on all remaining samples. Experimental results demonstrate that MoSo effectively mitigates severe performance degradation at high pruning ratios and achieves satisfactory performance across various settings. Existing approaches can be broadly categorized into three major groups: pruning by importance criteria [11, 42, 28, 27, 26, 46] , coverage or diversity-driven methods [48, 29, 33, 15] , and optimizationbased methods [38, 8, 23, 24, 21, 44] . Among these, the first group of studies is the most effective and popular. These studies assume that hard samples are critical and informative core-set samples, and thus, they design difficulty-based metrics to assess sample importance. Such metrics include prediction entropy [11] , forgetting [28] or memorization [46] score, gradient norm [27], E2LN (variance of prediction) [27], self-supervised prototype distance [42], diverse ensembles [26], and others.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper36
- Data-efficient Fine-tuning for LLM-based RecommendationXinyu Lin, Wenjie Wang, Yongqi Li, Shuo Yang 等SIGIR 2024 · 被引用 152 次
- Mind the Boundary: Coreset Selection via Reconstructing the Decision BoundaryShuo Yang, Zhe Cao, Sheng Guo, Ruiheng Zhang 等ICML 2024 · 被引用 25 次
- On the Diversity and Realism of Distilled Dataset: An Efficient Dataset Distillation ParadigmPeng Sun, Bei Shi, Daiwei Yu, Tao LinCVPR 2024 · 被引用 17 次
- How Far are AI-Generated Videos from Simulating the 3D Visual World: A Learned 3D Evaluation ApproachChirui Chang, Jiahui Liu, Zhengzhe Liu, Xiaoyang Lyu 等ICCV 2025 · 被引用 15 次
- Spanning Training Progress: Temporal Dual-Depth Scoring (TDDS) for Enhanced Dataset PruningXin Zhang, Jiawei Du, Yunsong Li, Weiying Xie 等CVPR 2024 · 被引用 12 次
它引用的顶会 Paper19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu 等NeurIPS 2022 · 被引用 2,727 次
- Beyond neural scaling laws: beating power law scaling via data pruningBen Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli 等NeurIPS 2022 · 被引用 720 次
相关 Paper
- Batch Loss Score for Dynamic Data PruningQing Zhou, Bingxuan Zhao, Tao Yang, Hongyuan Zhang 等CVPR 2026
- Dataset Pruning: Reducing Training Data by Examining Generalization InfluenceShuo Yang, Zeke Xie, Hanyu Peng, Min Xu 等ICLR 2023 · 被引用 21 次
- UNSEEN: Enhancing Dataset Pruning from a Generalization PerspectiveFurui Xu, Shaobo Wang, Jiajun Zhang, Chenghao Sun 等AAAI 2026
- Selecting Samples on Graphs: A Unified Dataset Pruning Framework for Lossless Training AccelerationDongyue Wu, Zilin Guo, Xiaoyu Li, Jiajia Liu 等ICML 2026
- Data Pruning by Information MaximizationHaoru Tan, Sitong Wu, Wei Huang, Shizhen Zhao 等ICLR 2025
