DynaMS: Dyanmic Margin Selection for Efficient Deep Learning
Jiaxing Wang, Yong Li, Jingwei Zhuo, Xupeng Shi, Weizhong Zhang, Lixing Gong, Tong Tao, Pengzhang Liu, Yongjun Bao, Weipeng Yan
Abstract
The great success of deep learning is largely driven by training over-parameterized models on massive datasets. To avoid excessive computation, extracting and training only on the most informative subset is drawing increasing attention. Nevertheless, it is still an open question how to select such a subset on which the model trained generalizes on par with the full data. In this paper, we propose dynamic margin selection (DynaMS). DynaMS leverages the distance from candidate samples to the classification boundary to construct the subset, and the subset is dynamically updated during model training. We show that DynaMS converges with large probability, and for the first time show both in theory and practice that dynamically updating the subset can result in better generalization. To reduce the additional computation incurred by the selection, a light parameter sharing proxy (PSP) is designed. PSP is able to faithfully evaluate instances following the underlying model, which is necessary for dynamic selection. Extensive analysis and experiments demonstrate the superiority of the proposed approach in data selection against many state-of-the-art counterparts on benchmark datasets.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 78c68181-e93d-44de-9e0e-b7638cdc0016Cited by top-tier papers2
- Making Scalable Meta Learning PracticalSang Keun Choe, Sanket Vaibhav Mehta, Hwijeen Ahn, Willie Neiswanger et al.NeurIPS 2023 · 28 citations
- Patch-Aware Sample Selection for Efficient Masked Image ModelingZhengyang Zhuge, Jiaxing Wang, Yong Li, Yongjun Bao et al.AAAI 2024 · 4 citations
Related papers
- Dataset Pruning: Reducing Training Data by Examining Generalization InfluenceShuo Yang, Zeke Xie, Hanyu Peng, Min Xu et al.ICLR 2023 · 21 citations
- Selection via Proxy: Efficient Data Selection for Deep LearningCody Coleman, Christopher Yeh, Stephen Mussmann, Baharan Mirzasoleiman et al.ICLR 2020 · 462 citations
- Efficient Representativeness-Aware Coreset SelectionZihao Cheng, Binrui Wu, Zhiwei Li, Yuesen Liao et al.NeurIPS 2025 · 1 citation
- Deep Active Learning by Leveraging Training DynamicsHaonan Wang, Wei Huang, Ziwei Wu, Hanghang Tong et al.NeurIPS 2022 · 49 citations
- Moderate Coreset: A Universal Method of Data Selection for Real-world Data-efficient Deep LearningXiaobo Xia, Jiale Liu, Jun Yu, Xu Shen et al.ICLR 2023
