CLIP the Gap: A Single Domain Generalization Approach for Object Detection
Vidit Vidit, Martin Engilberge, Mathieu Salzmann
Abstract
Single Domain Generalization (SDG) tackles the problem of training a model on a single source domain so that it generalizes to any unseen target domain. While this has been well studied for image classification, the literature on SDG object detection remains almost non-existent. To address the challenges of simultaneously learning robust object localization and representation, we propose to leverage a pre-trained vision-language model to introduce semantic domain concepts via textual prompts. We achieve this via a semantic augmentation strategy acting on the features extracted by the detector backbone, as well as a text-based classification loss. Our experiments evidence the benefits of our approach, outperforming by 10% the only existing SDG object detection method, Single-DGOD [52], on their own diverse weather-driving benchmark.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1aada8ec-3954-4d77-a2c6-b7c7e7f530ecCited by top-tier papers43
- PØDA: Prompt-driven Zero-shot Domain AdaptationMohammad Fahes, Tuan-Hung Vu, Andrei Bursuc, Patrick Pérez et al.ICCV 2023 · 82 citations
- Object-Aware Domain Generalization for Object DetectionWooju Lee, Dasol Hong, Hyungtae Lim, Hyun MyungAAAI 2024 · 58 citations
- Learning Domain-Aware Detection Head with Prompt TuningHaochen Li, Rui Zhang, Hantao Yao, Xinkai Song et al.NeurIPS 2023 · 40 citations
- Unbiased Faster R-CNN for Single-source Domain Generalized Object DetectionYajing Liu, Shijun Zhou, Xiyao Liu, Chunhui Hao et al.CVPR 2024 · 35 citations
- G-NAS: Generalizable Neural Architecture Search for Single Domain Generalization Object DetectionFan Wu, Jinling Gao, Lanqing Hong, Xinbing Wang et al.AAAI 2024 · 31 citations
Builds on25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 1,274 citations
- Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-TrainingGen Li, Nan Duan, Yuejian Fang, Ming Gong et al.AAAI 2020 · 966 citations
- Episodic Training for Domain GeneralizationDa Li, Jianshu Zhang, Yongxin Yang, Cong Liu et al.ICCV 2019 · 488 citations
Related papers
- Boosting Single-Domain Generalized Object Detection via Vision-Language Knowledge InteractionXiaoran Xu, Jiangang Yang, Wenyue Chong, Wenhui Shi et al.ACM MM 2025 · 2 citations
- PhysAug: A Physical-guided and Frequency-based Data Augmentation for Single-Domain Generalized Object DetectionXiaoran Xu, Jiangang Yang, Wenhui Shi, Siyuan Ding et al.AAAI 2025 · 15 citations
- Open-Vocabulary Domain Generalization in Urban-Scene SegmentationDong Zhao, Qi Zang, Nan Pu, Wenjing Li et al.CVPR 2026 · 3 citations
- Style Evolving along Chain-of-Thought for Unknown-Domain Object DetectionZihao Zhang, Aming Wu, Yahong HanCVPR 2025
- Adversarial Domain Prompt Tuning and Generation for Single Domain GeneralizationZhipeng Xu, De Cheng, Xinyang Jiang, Nannan Wang et al.CVPR 2025
