ORIENT: Submodular Mutual Information Measures for Data Subset Selection under Distribution Shift
Athresh Karanam, KrishnaTeja Killamsetty, Harsha Kokel, Rishabh K. Iyer
Abstract
Real-world machine-learning applications require robust models that generalize well to distribution shift settings, which is typical in real-world situations. Domain adaptation techniques aim to address this issue of distribution shift by minimizing the disparities between domains to ensure that the model trained on the source domain performs well on the target domain. Nevertheless, the existing domain adaptation methods are computationally very expensive. In this work, we aim to improve the efficiency of existing supervised domain adaptation (SDA) methods by using a subset of source data that is similar to target data for faster model training. Specifically, we propose ORIENT, a subset selection framework that uses the submodular mutual information (SMI) functions to select a source data subset similar to the target data for faster training. Additionally, we demonstrate how existing robust subset selection strategies, such as GLISTER, GRADMATCH, and CRAIG, when used with a held-out query set, fit within our proposed framework and demonstrate the connections with them. Finally, we empirically demonstrate that SDA approaches like d-SNE, CCSA, and standard Cross-entropy training, when employed together with ORIENT, achieve a) faster training and b) better performance on the target data. * equal contribution 36th Conference on Neural Information Processing Systems (NeurIPS 2022).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- Looking Beyond the Known: Towards a Data Discovery Guided Open-World Object DetectionAnay Majee, Amitesh Gangrade, Rishabh IyerNeurIPS 2025 · 5 citations
- Scalable Batch-Mode Deep Bayesian Active Learning via Equivalence Class AnnealingRenyu Zhang, Aly A. Khan, Robert L. Grossman, Yuxin ChenICLR 2023 · 1 citation
- Boosting the Uniqueness of Neural Networks Fingerprints with Informative TriggersZhuomeng Zhang, Fangqi Li, Hanyi Wang, Shi-Lin WangNeurIPS 2025 · 1 citation
- On the Support Vector Effect in DNNs: Rethinking Data Selection and AttributionSyed Hasan Amin Mahmood, Ming Yin, Rajiv KhannaKDD 2025
- Combatting Dimensional Collapse in LLM Pre-Training Data via Submodular File SelectionZiqing Fan, Siyuan Du, Shengchao Hu, Pingjie Wang et al.ICLR 2025
Builds on11
- Deep Batch Active Learning by Diverse, Uncertain Gradient Lower BoundsJordan T. Ash, Chicheng Zhang, Akshay Krishnamurthy, John Langford et al.ICLR 2020 · 974 citations
- Semi-Supervised Domain Adaptation via Minimax EntropyKuniaki Saito, Donghyun Kim, Stan Sclaroff, Trevor Darrell et al.ICCV 2019 · 725 citations
- Coresets for Data-efficient Training of Machine Learning ModelsBaharan Mirzasoleiman, Jeff A. Bilmes, Jure LeskovecICML 2020 · 494 citations
- GRAD-MATCH: Gradient Matching based Data Subset Selection for Efficient Deep Model TrainingKrishnaTeja Killamsetty, Durga Sivasubramanian, Ganesh Ramakrishnan, Abir De et al.ICML 2021 · 305 citations
- GLISTER: Generalization based Data Subset Selection for Efficient and Robust LearningKrishnaTeja Killamsetty, Durga Sivasubramanian, Ganesh Ramakrishnan, Rishabh K. IyerAAAI 2021 · 300 citations
Related papers
- Graph Data Selection for Domain Adaptation: A Model-Free ApproachTing-Wei Li, Ruizhong Qiu, Hanghang TongNeurIPS 2025 · 7 citations
- S3VAADA: Submodular Subset Selection for Virtual Adversarial Active Domain AdaptationHarsh Rangwani, Arihant Jain, Sumukh K. Aithal, R. Venkatesh BabuICCV 2021 · 33 citations
- RETRIEVE: Coreset Selection for Efficient and Robust Semi-Supervised LearningKrishnaTeja Killamsetty, Xujiang Zhao, Feng Chen, Rishabh K. IyerNeurIPS 2021 · 115 citations
- Semi-Supervised Domain Adaptation with Source Label AdaptationYu-Chu Yu, Hsuan-Tien LinCVPR 2023
- Project and Probe: Sample-Efficient Adaptation by Interpolating Orthogonal FeaturesAnnie S. Chen, Yoonho Lee, Amrith Setlur, Sergey Levine et al.ICLR 2024 · 5 citations
