Moderate Coreset: A Universal Method of Data Selection for Real-world Data-efficient Deep Learning
Xiaobo Xia, Jiale Liu, Jun Yu, Xu Shen, Bo Han, Tongliang Liu
摘要
Deep learning methods nowadays rely on massive data, resulting in substantial costs of data storage and model training. Data selection is a useful tool to alleviate such costs, where a coreset of massive data is extracted to practically perform on par with full data. Based on carefully-designed score criteria, existing methods first count the score of each data point and then select the data points whose scores lie in a certain range to construct a coreset. These methods work well in their respective preconceived scenarios but are not robust to the change of scenarios, since the optimal range of scores varies as the scenario changes. The issue limits the application of these methods, because realistic scenarios often mismatch preconceived ones, and it is inconvenient or unfeasible to tune the criteria and methods accordingly. In this paper, to address the issue, a concept of the moderate coreset is discussed. Specifically, given any score criterion of data selection, different scenarios prefer data points with scores in different intervals. As the score median is a proxy of the score distribution in statistics, the data points with scores close to the score median can be seen as a proxy of full data and generalize different scenarios, which are used to construct the moderate coreset. As a proof-of-concept, a universal method that inherits the moderate coreset and uses the distance of a data point to its class center as the score criterion, is proposed to meet complex realistic scenarios. Extensive experiments confirm the advance of our method over prior state-of-the-art methods, leading to a strong baseline for future research. The implementation is available at https://github.com/tmllab/Moderate-DS.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper72
- LESS: Selecting Influential Data for Targeted Instruction TuningMengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora 等ICML 2024 · 被引用 460 次
- HumanMAC: Masked Motion Completion for Human Motion PredictionLing-Hao Chen, Jiawei Zhang, Yewen Li, Yiren Pang 等ICCV 2023 · 被引用 106 次
- Data Pruning via Moving-one-Sample-outHaoru Tan, Sitong Wu, Fei Du, Yukang Chen 等NeurIPS 2023 · 被引用 91 次
- M3D: Dataset Condensation by Minimizing Maximum Mean DiscrepancyHansong Zhang, Shikun Li, Pengju Wang, Dan Zeng 等AAAI 2024 · 被引用 63 次
- Out-of-distribution Detection Learning with Unreliable Out-of-distribution SourcesHaotian Zheng, Qizhou Wang, Zhen Fang, Xiaobo Xia 等NeurIPS 2023 · 被引用 53 次
相关 Paper
- Refined Coreset Selection: Towards Minimal Coreset Size under Model Performance ConstraintsXiaobo Xia, Jiale Liu, Shaokun Zhang, Qingyun Wu 等ICML 2024 · 被引用 17 次
- Efficient Coreset Selection with Cluster-based MethodsChengliang Chai, Jiayi Wang, Nan Tang, Ye Yuan 等KDD 2023 · 被引用 19 次
- Mind the Boundary: Coreset Selection via Reconstructing the Decision BoundaryShuo Yang, Zhe Cao, Sheng Guo, Ruiheng Zhang 等ICML 2024 · 被引用 25 次
- Adaptive Dataset QuantizationMuquan Li, Dongyang Zhang, Qiang Dong, Xiurui Xie 等AAAI 2025 · 被引用 9 次
- Feature Distribution Matching by Optimal Transport for Effective and Robust Coreset SelectionWeiwei Xiao, Yongyong Chen, Qiben Shan, Yaowei Wang 等AAAI 2024 · 被引用 14 次
