Multimodal Unsupervised Domain Generalization by Retrieving Across the Modality Gap
Christopher Liao, Christian So, Theodoros Tsiligkaridis, Brian Kulis
摘要
Domain generalization (DG) is an important problem that involves learning a model which generalizes to unseen test domains by leveraging one or more source domains, under the assumption of shared label spaces. However, most DG methods assume access to abundant source data in the target label space, a requirement that proves overly stringent for numerous real-world applications, where acquiring the same label space as the target task is prohibitively expensive. For this setting, we tackle the multimodal version of the unsupervised domain generalization (MUDG) problem, which uses a large task-agnostic unlabeled source dataset during finetuning. Our framework relies only on the premise that the source dataset can be accurately and efficiently searched in a joint vision-language space. We make three contributions in the MUDG setting. Firstly, we show theoretically that cross-modal approximate nearest neighbor search suffers from low recall due to the large distance between text queries and the image centroids used for coarse quantization. Accordingly, we propose paired k-means, a simple clustering algorithm that improves nearest neighbor recall by storing centroids in query space instead of image space. Secondly, we propose an adaptive text augmentation scheme for target labels designed to improve zero-shot accuracy and diversify retrieved image data. Lastly, we present two simple but effective components to further improve downstream target accuracy. We compare against state-of-the-art name-only transfer, source-free DG and zero-shot (ZS) methods on their respective benchmarks and show consistent improvement in accuracy on 20 diverse datasets. Code is available:
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Is the Modality Gap a Bug or a Feature? A Robustness PerspectiveRhea Chowers, Oshri Naparstek, Udi Barzelay, Yair WeissCVPR 2026 · 被引用 4 次
- The Inter-Intra Modal Measure: A Predictive Lens on Fine-Tuning Outcomes in Vision-Language ModelsLaura Niss, Kevin Vogt-Lowell, Theodoros TsiligkaridisICCV 2025 · 被引用 1 次
- Toward Structural Multimodal Representations: Specialization, Selection, and Sparsification via Mixture-of-ExpertsHahyeon Choi, NOJUN KWAKICML 2026
它引用的顶会 Paper34
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna 等NeurIPS 2020 · 被引用 7,049 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath 等ICCV 2021 · 被引用 2,294 次
相关 Paper
- Promoting Semantic Connectivity: Dual Nearest Neighbors Contrastive Learning for Unsupervised Domain GeneralizationYuchen Liu, Yaoming Wang, Yabo Chen, Wenrui Dai 等CVPR 2023
- Retrieval Across Any Domains via Large-scale Pre-trained ModelJiexi Yan, Zhihui Yin, Chenghao Xu, Cheng Deng 等ICML 2024
- Progressive Distribution Bridging: Unsupervised Adaptation for Large-Scale Pre-Trained Models via Adaptive Auxiliary DataWeinan He, Yixin Zhang, Zilei WangICCV 2025 · 被引用 1 次
- Bridging Domain Generalization to Multimodal Domain Generalization via Unified RepresentationsHai Huang, Yan Xia, Sashuai Zhou, Hanting Wang 等ICCV 2025 · 被引用 2 次
- Context-Aware Multimodal PretrainingKarsten Roth, Zeynep Akata, Dima Damen, Ivana Balazevic 等CVPR 2025
