ORIENT: Submodular Mutual Information Measures for Data Subset Selection under Distribution Shift
Athresh Karanam, KrishnaTeja Killamsetty, Harsha Kokel, Rishabh K. Iyer
摘要
Real-world machine-learning applications require robust models that generalize well to distribution shift settings, which is typical in real-world situations. Domain adaptation techniques aim to address this issue of distribution shift by minimizing the disparities between domains to ensure that the model trained on the source domain performs well on the target domain. Nevertheless, the existing domain adaptation methods are computationally very expensive. In this work, we aim to improve the efficiency of existing supervised domain adaptation (SDA) methods by using a subset of source data that is similar to target data for faster model training. Specifically, we propose ORIENT, a subset selection framework that uses the submodular mutual information (SMI) functions to select a source data subset similar to the target data for faster training. Additionally, we demonstrate how existing robust subset selection strategies, such as GLISTER, GRADMATCH, and CRAIG, when used with a held-out query set, fit within our proposed framework and demonstrate the connections with them. Finally, we empirically demonstrate that SDA approaches like d-SNE, CCSA, and standard Cross-entropy training, when employed together with ORIENT, achieve a) faster training and b) better performance on the target data. * equal contribution 36th Conference on Neural Information Processing Systems (NeurIPS 2022).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Looking Beyond the Known: Towards a Data Discovery Guided Open-World Object DetectionAnay Majee, Amitesh Gangrade, Rishabh IyerNeurIPS 2025 · 被引用 5 次
- Scalable Batch-Mode Deep Bayesian Active Learning via Equivalence Class AnnealingRenyu Zhang, Aly A. Khan, Robert L. Grossman, Yuxin ChenICLR 2023 · 被引用 1 次
- Boosting the Uniqueness of Neural Networks Fingerprints with Informative TriggersZhuomeng Zhang, Fangqi Li, Hanyi Wang, Shi-Lin WangNeurIPS 2025 · 被引用 1 次
- On the Support Vector Effect in DNNs: Rethinking Data Selection and AttributionSyed Hasan Amin Mahmood, Ming Yin, Rajiv KhannaKDD 2025
- Combatting Dimensional Collapse in LLM Pre-Training Data via Submodular File SelectionZiqing Fan, Siyuan Du, Shengchao Hu, Pingjie Wang 等ICLR 2025
它引用的顶会 Paper11
- Deep Batch Active Learning by Diverse, Uncertain Gradient Lower BoundsJordan T. Ash, Chicheng Zhang, Akshay Krishnamurthy, John Langford 等ICLR 2020 · 被引用 974 次
- Semi-Supervised Domain Adaptation via Minimax EntropyKuniaki Saito, Donghyun Kim, Stan Sclaroff, Trevor Darrell 等ICCV 2019 · 被引用 725 次
- Coresets for Data-efficient Training of Machine Learning ModelsBaharan Mirzasoleiman, Jeff A. Bilmes, Jure LeskovecICML 2020 · 被引用 494 次
- GRAD-MATCH: Gradient Matching based Data Subset Selection for Efficient Deep Model TrainingKrishnaTeja Killamsetty, Durga Sivasubramanian, Ganesh Ramakrishnan, Abir De 等ICML 2021 · 被引用 305 次
- GLISTER: Generalization based Data Subset Selection for Efficient and Robust LearningKrishnaTeja Killamsetty, Durga Sivasubramanian, Ganesh Ramakrishnan, Rishabh K. IyerAAAI 2021 · 被引用 300 次
相关 Paper
- Graph Data Selection for Domain Adaptation: A Model-Free ApproachTing-Wei Li, Ruizhong Qiu, Hanghang TongNeurIPS 2025 · 被引用 7 次
- S3VAADA: Submodular Subset Selection for Virtual Adversarial Active Domain AdaptationHarsh Rangwani, Arihant Jain, Sumukh K. Aithal, R. Venkatesh BabuICCV 2021 · 被引用 33 次
- RETRIEVE: Coreset Selection for Efficient and Robust Semi-Supervised LearningKrishnaTeja Killamsetty, Xujiang Zhao, Feng Chen, Rishabh K. IyerNeurIPS 2021 · 被引用 115 次
- Semi-Supervised Domain Adaptation with Source Label AdaptationYu-Chu Yu, Hsuan-Tien LinCVPR 2023
- Project and Probe: Sample-Efficient Adaptation by Interpolating Orthogonal FeaturesAnnie S. Chen, Yoonho Lee, Amrith Setlur, Sergey Levine 等ICLR 2024 · 被引用 5 次
