Textual and Visual Guided Task Adaptation for Source-Free Cross-Domain Few-Shot Segmentation
Jianming Liu, Wenlong Qiu, Haitao Wei
Abstract
Few-Shot Segmentation(FSS) aims to efficient segmentation of new objects with few labeled samples. However, its performance significantly degrades when domain discrepancies exist between training and deployment. Cross-Domain Few-Shot Segmentation(CD-FSS) is proposed to mitigate such performance degradation. Current CD-FSS methods primarily sought to develop segmentation models on a source domain capable of cross-domain generalization. However, driven by escalating concerns over data privacy and the imperative to minimize data transfer and training expenses, the development of source-free CD-FSS approaches has become essential. In this work, we propose a source-free CD-FSS method that leverages both textual and visual information to facilitate target domain task adaptation without requiring source domain data. Specifically, we first append Task-Specific Attention Adapters (TSAA) to the feature pyramid of a pretrained backbone, which adapt multi-level features extracted from the shared pre-trained backbone to the target task. Then, the parameters of the TSAA are trained through a Visual-Visual Embedding Alignment (VVEA) module and a Text-Visual Embedding Alignment (TVEA) module. The VVEA module utilizes global-local visual features to align image features across different views, while the TVEA module leverages textual priors from pre-aligned multi-modal features (e.g., from CLIP) to guide cross-modal adaptation. By combining the outputs of these modules through dense comparison operations and subsequent fusion via skip connections, our method produces refined prediction masks. Under both 1-shot and 5-shot settings, the proposed approach achieves average segmentation accuracy improvements of 2.18% and 4.11%, respectively, across four cross-domain datasets, significantly outperforming state-of-the-art CD-FSS methods. Code are available at https://github.com/ljm198134/TVGTANet.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 830bccaf-6aa6-4b8f-a45b-52ae3de3062fCited by top-tier papers1
Ask how each one uses itBuilds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- PANet: Few-Shot Image Semantic Segmentation With Prototype AlignmentKaixin Wang, Jun Hao Liew, Yingtian Zou, Daquan Zhou et al.ICCV 2019 · 1,404 citations
- DenseCLIP: Language-Guided Dense Prediction with Context-Aware PromptingYongming Rao, Wenliang Zhao, Guangyi Chen, Yansong Tang et al.CVPR 2022 · 527 citations
- Image Segmentation Using Text and Image PromptsTimo Lüddecke, Alexander S. EckerCVPR 2022 · 457 citations
- Learning What Not to Segment: A New Perspective on Few-Shot SegmentationChunbo Lang, Gong Cheng, Binfei Tu, Junwei HanCVPR 2022 · 289 citations
Related papers
- Adapt Before Comparison: A New Perspective on Cross-Domain Few-Shot SegmentationJonas HerzogCVPR 2024
- Rethinking Prior Information Generation with CLIP for Few-Shot SegmentationJin Wang, Bingfeng Zhang, Jian Pang, Honglong Chen et al.CVPR 2024 · 27 citations
- Self-Disentanglement and Re-Composition for Cross-Domain Few-Shot SegmentationJintao Tong, Yixiong Zou, Guangyao Chen, Yuhua Li et al.ICML 2025
- Adapting In-Domain Few-Shot Segmentation to New Domains Without Source Domain RetrainingQi Fan, Kaiqi Liu, Nian Liu, Hisham Cholakkal et al.ICCV 2025 · 4 citations
- Adapter Naturally Serves as Decoupler for Cross-Domain Few-Shot Semantic SegmentationJintao Tong, Ran Ma, Yixiong Zou, Guangyao Chen et al.ICML 2025
