The Finer the Better: Towards Granular-aware Open-set Domain Generalization
Yunyun Wang, Zheng Duan, Xinyue Liao, Ke-Jia Chen, Songcan Chen
Abstract
Open-Set Domain Generalization (OSDG) aims to generalize over unseen target domains containing open classes, and the core challenge lies in identifying unknown samples never encountered during training. Recently, CLIP has exhibited impressive performance in OSDG, while it still falls into the dilemma between structural risk of known classes and open space risk from unknown classes, and easily suffers from over-confidence, especially when distinguishing known-like unknown samples. To this end, we propose a Semantic-enhanced CLIP (SeeCLIP) framework that leverages fine-grained semantics to boost unknown detection, so as to accommodate both risks and enable precise discrimination among categories. In SeeCLIP, we propose a semantic-aware prompt enhancement module to extract fine-grained key semantic features, and establish a fine-grained vision-language alignment. Duplex contrastive learning is proposed for prompt learning, which jointly optimizes duplex losses such that the unknown prompt is similar to known prompts, yet exhibits key semantic differences. We also design a semantic-guided diffusion module to enable nuanced capture in generation. By injecting perturbed key semantics into a diffusion model as control conditions, it generates the closest unknowns or pseudo-open samples with high similarity yet low belongingness to known classes. We formulate a generalization bound for OSDG, and show that SeeCLIP can achieve a lower generalization risk. Extensive experiments on benchmark datasets validate the superiority of SeeCLIP, it outperforms the SOTA methods by nearly 3% on accuracy and 5% on H-index, respectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Generalized Category DiscoverySagar Vaze, Kai Han, Andrea Vedaldi, Andrew ZissermanCVPR 2022 · 194 citations
- Generalizable Decision Boundaries: Dualistic Meta-Learning for Open Set Domain GeneralizationXiran Wang, Jian Zhang, Lei Qi, Yinghuan ShiICCV 2023 · 39 citations
Related papers
- Unknown Prompt, the only Lacuna: Unveiling CLIP's Potential for Open Domain GeneralizationMainak Singha, Ankit Jha, Shirsha Bose, Ashwin R. Nair et al.CVPR 2024 · 13 citations
- OSLoPrompt: Bridging Low-Supervision Challenges and Open-Set Domain Generalization in CLIPMohamad Hassan N C, Divyam Gupta, Mainak Singha, Sai Bhargav Rongali et al.CVPR 2025
- Open-Set Fine-Grained Retrieval via Prompting Vision-Language EvaluatorShijie Wang, Jianlong Chang, Haojie Li, Zhihui Wang et al.CVPR 2023
- CLIP-driven Coarse-to-fine Semantic Guidance for Fine-grained Open-set Semi-supervised LearningXiaokun Li, Yaping Huang, Qingji GuanCVPR 2025
- DPSeg: Dual-Prompt Cost Volume Learning for Open-Vocabulary Semantic SegmentationZiyu Zhao, Xiaoguang Li, Lingjia Shi, Nasrin Imanpour et al.CVPR 2025
