Efficient Data Subset Selection to Generalize Training Across Models: Transductive and Inductive Networks
Eeshaan Jain, Tushar Nandy, Gaurav Aggarwal, Ashish Tendulkar, Rishabh K. Iyer, Abir De
Abstract
Existing subset selection methods for efficient learning predominantly employ discrete combinatorial and model-specific approaches which lack generalizability. For an unseen architecture, one cannot use the subset chosen for a different model. To tackle this problem, we propose , a trainable subset selection framework, that generalizes across architectures. Here, we first introduce an attention-based neural gadget that leverages the graph structure of architectures and acts as a surrogate to trained deep neural networks for quick model prediction. Then, we use these predictions to build subset samplers. This naturally provides us two variants of . The first variant is transductive (called as Transductive-) which computes the subset separately for each model by solving a small optimization problem. Such an optimization is still super fast, thanks to the replacement of explicit model training by the model approximator. The second variant is inductive (called as Inductive-) which computes the subset using a trained subset selector, without any optimization. Our experiments show that our model outperforms several methods across several real datasets
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fe0e39d0-4ed9-4224-9b74-2528c1f6fac5Cited by top-tier papers6
- "What Data Benefits My Classifier?" Enhancing Model Performance and Interpretability through Influence-Based Data SelectionAnshuman Chhabra, Peizhao Li, Prasant Mohapatra, Hongfu LiuICLR 2024 · 32 citations
- SCoRe: Submodular Combinatorial Representation LearningAnay Majee, Suraj Kothawade, Krishnateja Killamsetty, Rishabh K. IyerICML 2024 · 7 citations
- Looking Beyond the Known: Towards a Data Discovery Guided Open-World Object DetectionAnay Majee, Amitesh Gangrade, Rishabh IyerNeurIPS 2025 · 5 citations
- Efficient Core-set Selection for Deep Learning Through Squared Loss MinimizationJianting ChenICML 2025
- Outlier Gradient Analysis: Efficiently Identifying Detrimental Training Samples for Deep Learning ModelsAnshuman Chhabra, Bo Li, Jian Chen, Prasant Mohapatra et al.ICML 2025
Builds on13
- NAS-Bench-201: Extending the Scope of Reproducible Neural Architecture SearchXuanyi Dong, Yi YangICLR 2020 · 825 citations
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 806 citations
- Beyond neural scaling laws: beating power law scaling via data pruningBen Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli et al.NeurIPS 2022 · 720 citations
- Coresets for Data-efficient Training of Machine Learning ModelsBaharan Mirzasoleiman, Jeff A. Bilmes, Jure LeskovecICML 2020 · 494 citations
- Are Transformers universal approximators of sequence-to-sequence functions?Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank J. Reddi et al.ICLR 2020 · 481 citations
Related papers
- Rapid Neural Architecture Search by Learning to Generate Graphs from DatasetsHayeon Lee, Eunyoung Hyung, Sung Ju HwangICLR 2021 · 57 citations
- Fitting the Search Space of Weight-sharing NAS with Graph Convolutional NetworksXin Chen, Lingxi Xie, Jun Wu, Longhui Wei et al.AAAI 2021 · 15 citations
- A Semi-Supervised Assessor of Neural ArchitecturesYehui Tang, Yunhe Wang, Yixing Xu, Hanting Chen et al.CVPR 2020
- Transfer NAS with Meta-learned Bayesian SurrogatesGresa Shala, Thomas Elsken, Frank Hutter, Josif GrabockaICLR 2023
- Efficient Subgraph GNNs by Learning Effective Selection PoliciesBeatrice Bevilacqua, Moshe Eliasof, Eli A. Meirom, Bruno Ribeiro et al.ICLR 2024 · 20 citations
