Cap2Seg: Inferring Semantic and Spatial Context from Captions for Zero-Shot Image Segmentation
Guiyu Tian, Shuai Wang, Jie Feng, Li Zhou, Yadong Mu
Abstract
Zero-shot image segmentation refers to the task of segmenting pixels from specific unseen semantic class. Previous methods mainly rely on historic segmentation tasks, such as using semantic embedding or word embedding of class names to infer a new segmentation model. In this work we describe Cap2Seg, a novel solution of zero-shot image segmentation that harnesses accompanying image captions for intelligently inferring spatial and semantic context for the zero-shot image segmentation task. As our main insight, image captions often implicitly entail the occurrence of a new class in an image and its most-confident spatial distribution. We define a contextual entailment question (CEQ) that tailors BERT-like text models. In specific, the proposed networks for inferring unseen classes consists of three branches (global / local / semi-global), which infer labels of unseen class from image level, image-stripe level or pixel level respectively. Comprehensive experiments and ablation studies are conducted on two image benchmarks, COCO-stuff and Pascal VOC. All clearly demonstrate the effectiveness of the proposed Cap2Seg, including a set of hardest unseen classes (i.e., image captions do not literally contain the class names and direct matching for inference fails).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext be59ad79-fa92-40ae-956c-1954daab5594Cited by top-tier papers4
- Decoupling Zero-Shot Semantic SegmentationJian Ding, Nan Xue, Gui-Song Xia, Dengxin DaiCVPR 2022 · 255 citations
- Perceptual Grouping in Contrastive Vision-Language ModelsKanchana Ranasinghe, Brandon McKinzie, Sachin Ravi, Yinfei Yang et al.ICCV 2023 · 88 citations
- Open-Vocabulary Instance Segmentation via Robust Cross-Modal Pseudo-LabelingDat Huynh, Jason Kuen, Zhe Lin, Jiuxiang Gu et al.CVPR 2022 · 78 citations
- Transferring CLIP's Knowledge into Zero-Shot Point Cloud Semantic SegmentationYuanbin Wang, Shaofei Huang, Yulu Gao, Zhen Wang et al.ACM MM 2023 · 17 citations
Builds on1
Related papers
- Context-aware Feature Generation For Zero-shot Semantic SegmentationZhangxuan Gu, Siyuan Zhou, Li Niu, Zihan Zhao et al.ACM MM 2020 · 111 citations
- Language-driven Semantic SegmentationBoyi Li, Kilian Q. Weinberger, Serge J. Belongie, Vladlen Koltun et al.ICLR 2022 · 885 citations
- Primitive Generation and Semantic-Related Alignment for Universal Zero-Shot SegmentationShuting He, Henghui Ding, Wei JiangCVPR 2023
- Exploring Regional Clues in CLIP for Zero-Shot Semantic SegmentationYi Zhang, Meng-Hao Guo, Miao Wang, Shi-Min HuCVPR 2024 · 20 citations
- Zero-shot Referring Image Segmentation with Global-Local Context FeaturesSeonghoon Yu, Paul Hongsuck Seo, Jeany SonCVPR 2023
