Efficient Data Subset Selection to Generalize Training Across Models: Transductive and Inductive Networks
Eeshaan Jain, Tushar Nandy, Gaurav Aggarwal, Ashish Tendulkar, Rishabh K. Iyer, Abir De
摘要
Existing subset selection methods for efficient learning predominantly employ discrete combinatorial and model-specific approaches which lack generalizability. For an unseen architecture, one cannot use the subset chosen for a different model. To tackle this problem, we propose , a trainable subset selection framework, that generalizes across architectures. Here, we first introduce an attention-based neural gadget that leverages the graph structure of architectures and acts as a surrogate to trained deep neural networks for quick model prediction. Then, we use these predictions to build subset samplers. This naturally provides us two variants of . The first variant is transductive (called as Transductive-) which computes the subset separately for each model by solving a small optimization problem. Such an optimization is still super fast, thanks to the replacement of explicit model training by the model approximator. The second variant is inductive (called as Inductive-) which computes the subset using a trained subset selector, without any optimization. Our experiments show that our model outperforms several methods across several real datasets
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- "What Data Benefits My Classifier?" Enhancing Model Performance and Interpretability through Influence-Based Data SelectionAnshuman Chhabra, Peizhao Li, Prasant Mohapatra, Hongfu LiuICLR 2024 · 被引用 32 次
- SCoRe: Submodular Combinatorial Representation LearningAnay Majee, Suraj Kothawade, Krishnateja Killamsetty, Rishabh K. IyerICML 2024 · 被引用 7 次
- Looking Beyond the Known: Towards a Data Discovery Guided Open-World Object DetectionAnay Majee, Amitesh Gangrade, Rishabh IyerNeurIPS 2025 · 被引用 5 次
- Efficient Core-set Selection for Deep Learning Through Squared Loss MinimizationJianting ChenICML 2025
- Outlier Gradient Analysis: Efficiently Identifying Detrimental Training Samples for Deep Learning ModelsAnshuman Chhabra, Bo Li, Jian Chen, Prasant Mohapatra 等ICML 2025
它引用的顶会 Paper13
- NAS-Bench-201: Extending the Scope of Reproducible Neural Architecture SearchXuanyi Dong, Yi YangICLR 2020 · 被引用 825 次
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 被引用 806 次
- Beyond neural scaling laws: beating power law scaling via data pruningBen Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli 等NeurIPS 2022 · 被引用 720 次
- Coresets for Data-efficient Training of Machine Learning ModelsBaharan Mirzasoleiman, Jeff A. Bilmes, Jure LeskovecICML 2020 · 被引用 494 次
- Are Transformers universal approximators of sequence-to-sequence functions?Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank J. Reddi 等ICLR 2020 · 被引用 481 次
相关 Paper
- Rapid Neural Architecture Search by Learning to Generate Graphs from DatasetsHayeon Lee, Eunyoung Hyung, Sung Ju HwangICLR 2021 · 被引用 57 次
- Fitting the Search Space of Weight-sharing NAS with Graph Convolutional NetworksXin Chen, Lingxi Xie, Jun Wu, Longhui Wei 等AAAI 2021 · 被引用 15 次
- A Semi-Supervised Assessor of Neural ArchitecturesYehui Tang, Yunhe Wang, Yixing Xu, Hanting Chen 等CVPR 2020
- Transfer NAS with Meta-learned Bayesian SurrogatesGresa Shala, Thomas Elsken, Frank Hutter, Josif GrabockaICLR 2023
- Efficient Subgraph GNNs by Learning Effective Selection PoliciesBeatrice Bevilacqua, Moshe Eliasof, Eli A. Meirom, Bruno Ribeiro 等ICLR 2024 · 被引用 20 次
