Data Pruning by Information Maximization
Haoru Tan, Sitong Wu, Wei Huang, Shizhen Zhao, Xiaojuan Qi
Abstract
In this paper, we present InfoMax, a novel data pruning method, also known as coreset selection, designed to maximize the information content of selected samples while minimizing redundancy. By doing so, InfoMax enhances the overall informativeness of the coreset. The information of individual samples is measured by importance scores, which capture their influence or difficulty in model learning. To quantify redundancy, we use pairwise sample similarities, based on the premise that similar samples contribute similarly to the learning process. We formalize the coreset selection problem as a discrete quadratic programming (DQP) task, with the objective of maximizing the total information content, represented as the sum of individual sample contributions minus the redundancies introduced by similar samples within the coreset. To ensure practical scalability, we introduce an efficient gradient-based solver, complemented by sparsification techniques applied to the similarity matrix and dataset partitioning strategies. This enables InfoMax to seamlessly scale to datasets with millions of samples. Extensive experiments demonstrate the superior performance of InfoMax in various data pruning tasks, including image classification, vision-language pre-training, and instruction tuning for large language models. Research in this field can be broadly divided into two categories: score-based methods (Sorscher et al., 2022; Radford et al., 2021) and geometrybased methods (Sener & Savarese, 2017; Ash et al., 2020) . Score-based methods focus on developing metrics to evaluate a sample's informativeness, such as prediction uncertainty (Har-Peled et al., 2007) , loss value (Cody Coleman et al., 2019) , or influence score (Xia et al., 2024; Tan et al., 2023) . However, as shown in Figure 2 (a), these methods often select samples densely concentrated in regions with the highest scores, leading to redundancies and failing to consider simpler samples with lower scores, which results in biased selections. Recent work (Sorscher et al., 2022) highlights that even simple samples are important for improving model generalization. On the other hand, ˚Equal Contribution. : Corresponding Author
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 896c814c-a69c-411b-b3d5-ae4d4051707aCited by top-tier papers11
- How Far are AI-Generated Videos from Simulating the 3D Visual World: A Learned 3D Evaluation ApproachChirui Chang, Jiahui Liu, Zhengzhe Liu, Xiaoyang Lyu et al.ICCV 2025 · 15 citations
- Towards More Diverse and Challenging Pre-Training for Point Cloud Learning: Self-Supervised Cross Reconstruction with Decoupled ViewsXiangdong Zhang, Shaofeng Zhang, Junchi YanICCV 2025 · 4 citations
- Equipping Vision Foundation Model with Mixture of Experts for Out-of-Distribution DetectionShizhen Zhao, Jiahui Liu, Xin Wen, Haoru Tan et al.ICCV 2025 · 3 citations
- GenHancer: Imperfect Generative Models are Secretly Strong Vision-Centric EnhancersShijie Ma, Yuying Ge, Teng Wang, Yuxin Guo et al.ICCV 2025 · 1 citation
- Unsupervised Process-Aware Coreset Selection for In-Context LearningWei Zheng, Zijie Wang, Xin Li, Bin Gong et al.ICML 2026
Builds on30
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
Related papers
- D2 Pruning: Message Passing for Balancing Diversity & Difficulty in Data PruningAdyasha Maharana, Prateek Yadav, Mohit BansalICLR 2024 · 19 citations
- Less is More: High-value Data Selection for Visual Instruction TuningZikang Liu, Kun Zhou, Wayne Xin Zhao, Dawei Gao et al.ACM MM 2025 · 3 citations
- Exploring Learning Complexity for Efficient Downstream Dataset PruningWenyu Jiang, Zhenlong Liu, Zejian Xie, Songxin Zhang et al.ICLR 2025
- InfoBatch: Lossless Training Speed Up by Unbiased Dynamic Data PruningZiheng Qin, Kai Wang, Zangwei Zheng, Jianyang Gu et al.ICLR 2024 · 94 citations
- STAFF: Speculative Coreset Selection for Task-Specific Fine-tuningXiaoyu Zhang, Juan Zhai, Shiqing Ma, Chao Shen et al.ICLR 2025
