Hierarchical Dataset Selection for High-Quality Data Sharing
Xiaona Zhou, Yingyan Zeng, Ran Jin, Ismini Lourentzou
摘要
The success of modern machine learning hinges on access to high-quality training data. In many real-world scenarios, such as acquiring data from public repositories or sharing across institutions, data is naturally organized into discrete datasets that vary in relevance, quality, and utility. Selecting which repositories or institutions to search for useful datasets, and which datasets to incorporate into model training are therefore critical decisions, yet most existing methods select individual samples and treat all data as equally relevant, ignoring differences between datasets and their sources. In this work, we formalize the task of dataset selection: selecting entire datasets from a large, heterogeneous pool to improve downstream performance under resource constraints. We propose Dataset Selection via Hierarchies (DaSH), a dataset selection method that models utility at both dataset and group (e.g., collections, institutions) levels, enabling efficient generalization from limited observations. Across two public benchmarks (Digit-Five and DomainNet), DaSH outperforms state-of-the-art data selection baselines by up to 26.2% in accuracy, while requiring significantly fewer exploration steps. Ablations show DaSH is robust to low-resource settings and lack of relevant datasets, making it suitable for scalable and adaptive dataset selection in practical multi-source learning workflows.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- Moment Matching for Multi-Source Domain AdaptationXingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang 等ICCV 2019 · 被引用 2,239 次
- Selection via Proxy: Efficient Data Selection for Deep LearningCody Coleman, Christopher Yeh, Stephen Mussmann, Baharan Mirzasoleiman 等ICLR 2020 · 被引用 462 次
- GLISTER: Generalization based Data Subset Selection for Efficient and Robust LearningKrishnaTeja Killamsetty, Durga Sivasubramanian, Ganesh Ramakrishnan, Rishabh K. IyerAAAI 2021 · 被引用 300 次
- Data Valuation using Reinforcement LearningJinsung Yoon, Sercan Ömer Arik, Tomas PfisterICML 2020 · 被引用 236 次
- CLDA: Contrastive Learning for Semi-Supervised Domain AdaptationAnkit SinghNeurIPS 2021 · 被引用 153 次
相关 Paper
- HCDS: Hierarchical Clustering for Cold-Start Few-Shot Data SelectionYuhua Zhao, Zhixin Han, Xunzhi Wang, Bitong Luo 等SIGIR 2025
- Learning a Universal Template for Few-shot Dataset GeneralizationEleni Triantafillou, Hugo Larochelle, Richard S. Zemel, Vincent DumoulinICML 2021 · 被引用 113 次
- Meta-Dataset: A Dataset of Datasets for Learning to Learn from Few ExamplesEleni Triantafillou, Tyler Zhu, Vincent Dumoulin, Pascal Lamblin 等ICLR 2020 · 被引用 692 次
- The More, The Better? Active Silencing of Non-Positive Transfer for Efficient Multi-Domain Few-Shot ClassificationXingxing Zhang, Zhizhe Liu, Weikai Yang, Liyuan Wang 等ACM MM 2022 · 被引用 1 次
- DsDm: Model-Aware Dataset Selection with DatamodelsLogan Engstrom, Axel Feldmann, Aleksander MadryICML 2024 · 被引用 105 次
