Beyond Efficiency: Molecular Data Pruning for Enhanced Generalization
Dingshuo Chen, Zhixun Li, Yuyan Ni, Guibin Zhang, Ding Wang, Qiang Liu, Shu Wu, Jeffrey Xu Yu, Liang Wang
Abstract
With the emergence of various molecular tasks and massive datasets, how to perform efficient training has become an urgent yet under-explored issue in the area. Data pruning (DP), as an oft-stated approach to saving training burdens, filters out less influential samples to form a coreset for training. However, the increasing reliance on pretrained models for molecular tasks renders traditional in-domain DP methods incompatible. Therefore, we propose a Molecular data Pruning framework for enhanced Generalization (MolPeg), which focuses on the source-free data pruning scenario, where data pruning is applied with pretrained models. By maintaining two models with different updating paces during training, we introduce a novel scoring function to measure the informativeness of samples based on the loss discrepancy. As a plug-and-play framework, MolPeg realizes the perception of both source and target domain and consistently outperforms existing DP methods across four downstream tasks. Remarkably, it can surpass the performance obtained from full-dataset training, even when pruning up to 60-70% of the data on HIV and PCBA dataset. Our work suggests that the discovery of effective data-pruning metrics could provide a viable path to both enhanced efficiency and superior generalization in transfer learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b63e4f4c-5741-4a65-9235-9dc4249ca4b1Cited by top-tier papers8
- GDeR: Safeguarding Efficiency, Balancing, and Robustness via Prototypical Graph PruningGuibin Zhang, Haonan Dong, Yuchen Zhang, Zhixun Li et al.NeurIPS 2024 · 9 citations
- Towards Objective Fine-tuning: How LLMs' Prior Knowledge Causes Potential Poor Calibration?Ziming Wang, Zeyu Shi, Haoyi Zhou, Shiqi Gao et al.ACL 2025 · 6 citations
- OrderDP: A Theoretically Guaranteed Lossless Dynamic Data Pruning FrameworkChenhan Jin, Shengze Xu, Qingsong Wang, Fan JIA et al.ICLR 2026 · 3 citations
- Learning Dynamics of Logits Debiasing for Long-Tailed Semi-Supervised LearningYue Cheng, Jiajun Zhang, Xiaohui Gao, Weiwei Xing et al.ICLR 2026 · 1 citation
- Efficient Core-set Selection for Deep Learning Through Squared Loss MinimizationJianting ChenICML 2025
Builds on26
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong et al.NeurIPS 2020 · 3,935 citations
- Graph Contrastive Learning with AugmentationsYuning You, Tianlong Chen, Yongduo Sui, Ting Chen et al.NeurIPS 2020 · 3,042 citations
- Strategies for Pre-training Graph Neural NetworksWeihua Hu, Bowen Liu, Joseph Gomes, Marinka Zitnik et al.ICLR 2020 · 1,744 citations
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 806 citations
- Equivariant message passing for the prediction of tensorial properties and molecular spectraKristof Schütt, Oliver T. Unke, Michael GasteggerICML 2021 · 736 citations
Related papers
- Selectivity Drives Productivity: Efficient Dataset Pruning for Enhanced Transfer LearningYihua Zhang, Yimeng Zhang, Aochuan Chen, Jinghan Jia et al.NeurIPS 2023 · 18 citations
- UNSEEN: Enhancing Dataset Pruning from a Generalization PerspectiveFurui Xu, Shaobo Wang, Jiajun Zhang, Chenghao Sun et al.AAAI 2026
- InfoBatch: Lossless Training Speed Up by Unbiased Dynamic Data PruningZiheng Qin, Kai Wang, Zangwei Zheng, Jianyang Gu et al.ICLR 2024 · 94 citations
- Advancing Graph Foundation Models: A Data-Centric PerspectiveYuhan Li, Yuyao Wang, Jianheng Tang, Heng Chang et al.KDD 2025
- Data Pruning via Moving-one-Sample-outHaoru Tan, Sitong Wu, Fei Du, Yukang Chen et al.NeurIPS 2023 · 91 citations
