CLIP the Gap: A Single Domain Generalization Approach for Object Detection
Vidit Vidit, Martin Engilberge, Mathieu Salzmann
摘要
Single Domain Generalization (SDG) tackles the problem of training a model on a single source domain so that it generalizes to any unseen target domain. While this has been well studied for image classification, the literature on SDG object detection remains almost non-existent. To address the challenges of simultaneously learning robust object localization and representation, we propose to leverage a pre-trained vision-language model to introduce semantic domain concepts via textual prompts. We achieve this via a semantic augmentation strategy acting on the features extracted by the detector backbone, as well as a text-based classification loss. Our experiments evidence the benefits of our approach, outperforming by 10% the only existing SDG object detection method, Single-DGOD [52], on their own diverse weather-driving benchmark.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper43
- PØDA: Prompt-driven Zero-shot Domain AdaptationMohammad Fahes, Tuan-Hung Vu, Andrei Bursuc, Patrick Pérez 等ICCV 2023 · 被引用 82 次
- Object-Aware Domain Generalization for Object DetectionWooju Lee, Dasol Hong, Hyungtae Lim, Hyun MyungAAAI 2024 · 被引用 58 次
- Learning Domain-Aware Detection Head with Prompt TuningHaochen Li, Rui Zhang, Hantao Yao, Xinkai Song 等NeurIPS 2023 · 被引用 40 次
- Unbiased Faster R-CNN for Single-source Domain Generalized Object DetectionYajing Liu, Shijun Zhou, Xiyao Liu, Chunhui Hao 等CVPR 2024 · 被引用 35 次
- G-NAS: Generalizable Neural Architecture Search for Single Domain Generalization Object DetectionFan Wu, Jinling Gao, Lanqing Hong, Xinbing Wang 等AAAI 2024 · 被引用 31 次
它引用的顶会 Paper25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 被引用 1,274 次
- Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-TrainingGen Li, Nan Duan, Yuejian Fang, Ming Gong 等AAAI 2020 · 被引用 966 次
- Episodic Training for Domain GeneralizationDa Li, Jianshu Zhang, Yongxin Yang, Cong Liu 等ICCV 2019 · 被引用 488 次
相关 Paper
- Boosting Single-Domain Generalized Object Detection via Vision-Language Knowledge InteractionXiaoran Xu, Jiangang Yang, Wenyue Chong, Wenhui Shi 等ACM MM 2025 · 被引用 2 次
- PhysAug: A Physical-guided and Frequency-based Data Augmentation for Single-Domain Generalized Object DetectionXiaoran Xu, Jiangang Yang, Wenhui Shi, Siyuan Ding 等AAAI 2025 · 被引用 15 次
- Open-Vocabulary Domain Generalization in Urban-Scene SegmentationDong Zhao, Qi Zang, Nan Pu, Wenjing Li 等CVPR 2026 · 被引用 3 次
- Style Evolving along Chain-of-Thought for Unknown-Domain Object DetectionZihao Zhang, Aming Wu, Yahong HanCVPR 2025
- Adversarial Domain Prompt Tuning and Generation for Single Domain GeneralizationZhipeng Xu, De Cheng, Xinyang Jiang, Nannan Wang 等CVPR 2025
