D2 Pruning: Message Passing for Balancing Diversity & Difficulty in Data Pruning
Adyasha Maharana, Prateek Yadav, Mohit Bansal
摘要
In recent years, data quality has emerged as an important factor for training massive models. Analytical theories suggest that higher-quality data can lead to lower test errors in models trained on a fixed data budget. Moreover, a model can be trained on a lower compute budget without compromising performance if a dataset can be stripped of its redundancies. Coreset selection (or data pruning) seeks to select a subset of the training data so as to maximize the performance of models trained on this subset, also referred to as coreset. There are two dominant approaches: (1) geometry-based data selection for maximizing data diversity in the coreset, and (2) functions that assign difficulty scores to samples based on training dynamics. Optimizing for data diversity leads to a coreset that is biased towards easier samples, whereas, selection by difficulty ranking omits easy samples that are necessary for the training of deep learning models. This demonstrates that data diversity and importance scores are two complementary factors that need to be jointly considered during coreset selection. In this work, we represent a dataset as an undirected graph and propose a novel pruning algorithm, D 2 PRUNING, that uses message passing over this dataset graph for coreset selection. D 2 PRUNING updates the difficulty scores of each example by incorporating the difficulty of its neighboring examples in the dataset graph. Then, these updated difficulty scores direct a graph-based sampling method to select a coreset that encapsulates both diverse and difficult regions of the dataset space. We evaluate supervised and self-supervised versions of our method on various vision and NLP datasets. Results show that D 2 PRUNING improves coreset selection over previous state-of-the-art methods at low-to-medium pruning rates. Additionally, we find that using D 2 PRUNING for filtering large multimodal datasets leads to increased diversity in the dataset and improved generalization of pretrained models. Our work shows that D 2 PRUNING is a versatile framework for understanding and processing datasets. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- CoIDO: Efficient Data Selection for Visual Instruction Tuning via Coupled Importance-Diversity OptimizationYichen Yan, Ming Zhong, Qi Zhu, Xiaoling Gu 等NeurIPS 2025 · 被引用 8 次
- Exploring Parameter-Efficient Fine-Tuning of Large Language Model on Automated Program RepairGuochang Li, Chen Zhi, Jialiang Chen, Junxiao Han 等ASE 2024 · 被引用 8 次
- AdaSTaR: Adaptive Data Sampling for Training Self-Taught ReasonersReiss Koh, Wonbeen Oh, Jaein Jang, Minhyung Lee 等NeurIPS 2025 · 被引用 8 次
- GORACS: Group-level Optimal Transport-guided Coreset Selection for LLM-based Recommender SystemsTiehua Mei, Hengrui Chen, Peng Yu, Jiaqing Liang 等KDD 2025 · 被引用 3 次
- FairDeDup: Detecting and Mitigating Vision-Language Fairness Disparities in Semantic Dataset DeduplicationEric Slyman, Stefan Lee, Scott Cohen, Kushal KafleCVPR 2024 · 被引用 3 次
它引用的顶会 Paper25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal 等NeurIPS 2020 · 被引用 5,249 次
- Directional Message Passing for Molecular GraphsJohannes Klicpera, Janek Groß, Stephan GünnemannICLR 2020 · 被引用 1,079 次
- Deep Batch Active Learning by Diverse, Uncertain Gradient Lower BoundsJordan T. Ash, Chicheng Zhang, Akshay Krishnamurthy, John Langford 等ICLR 2020 · 被引用 974 次
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 被引用 806 次
相关 Paper
- Data Pruning by Information MaximizationHaoru Tan, Sitong Wu, Wei Huang, Shizhen Zhao 等ICLR 2025
- Spanning Training Progress: Temporal Dual-Depth Scoring (TDDS) for Enhanced Dataset PruningXin Zhang, Jiawei Du, Yunsong Li, Weiying Xie 等CVPR 2024 · 被引用 12 次
- ELFS: Label-Free Coreset Selection with Proxy Training DynamicsHaizhong Zheng, Elisa Tsai, Yifu Lu, Jiachen Sun 等ICLR 2025
- UNSEEN: Enhancing Dataset Pruning from a Generalization PerspectiveFurui Xu, Shaobo Wang, Jiajun Zhang, Chenghao Sun 等AAAI 2026
- GDeR: Safeguarding Efficiency, Balancing, and Robustness via Prototypical Graph PruningGuibin Zhang, Haonan Dong, Yuchen Zhang, Zhixun Li 等NeurIPS 2024 · 被引用 9 次
