Betrayed by Captions: Joint Caption Grounding and Generation for Open Vocabulary Instance Segmentation
Jianzong Wu, Xiangtai Li, Henghui Ding, Xia Li, Guangliang Cheng, Yunhai Tong, Chen Change Loy
Abstract
In this work, we focus on open vocabulary instance segmentation to expand a segmentation model to classify and segment instance-level novel categories. Previous approaches have relied on massive caption datasets and complex pipelines to establish one-to-one mappings between image regions and words in captions. However, such methods build noisy supervision by matching non-visible words to image regions, such as adjectives and verbs. Meanwhile, context words are also important for inferring the existence of novel objects as they show high inter-correlations with novel categories. To overcome these limitations, we devise a joint Caption Grounding and Generation (CGG) framework, which incorporates a novel grounding loss that only focuses on matching object nouns to improve learning efficiency. We also introduce a caption generation head that enables additional supervision and contextual modeling as a complementation to the grounding loss. Our analysis and results demonstrate that grounding and generation components complement each other, significantly enhancing the segmentation performance for novel classes. Experiments on the COCO dataset with two settings: Open Vocabulary Instance Segmentation (OVIS) and Open Set Panoptic Segmentation (OSPS) demonstrate the superiority of the CGG. Specifically, CGG achieves a substantial improvement of 6.8% mAP for novel classes without extra data on the OVIS task and 15% PQ improvements for novel classes on the OSPS benchmark.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f9a5eadb-55e8-486c-b0a3-bfddf1b81971Cited by top-tier papers10
- Open3DIS: Open-Vocabulary 3D Instance Segmentation with 2D Mask GuidancePhuc D. A. Nguyen, Tuan Duc Ngo, Evangelos Kalogerakis, Chuang Gan et al.CVPR 2024 · 45 citations
- MaskClustering: View Consensus Based Mask Graph Clustering for Open-Vocabulary 3D Instance SegmentationMi Yan, Jiazhao Zhang, Yan Zhu, He WangCVPR 2024 · 22 citations
- Towards Language-Driven Video Inpainting via Multimodal Large Language ModelsJianzong Wu, Xiangtai Li, Chenyang Si, Shangchen Zhou et al.CVPR 2024 · 20 citations
- MV3DIS: Multi-View Mask Matching via 3D Guides for Zero-Shot 3D Instance SegmentationYibo Zhao, Yigong Zhang, Jin XieCVPR 2026 · 1 citation
- Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional VideosSagnik Majumder, Tushar Nagarajan, Ziad Al-Halah, Reina Pradhan et al.CVPR 2025
Builds on29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- YOLACT: Real-Time Instance SegmentationDaniel Bolya, Chong Zhou, Fanyi Xiao, Yong Jae LeeICCV 2019 · 2,075 citations
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 1,274 citations
- SOLOv2: Dynamic and Fast Instance SegmentationXinlong Wang, Rufeng Zhang, Tao Kong, Lei Li et al.NeurIPS 2020 · 1,193 citations
Related papers
- Mask-Free OVIS: Open-Vocabulary Instance Segmentation without Manual Mask AnnotationsVibashan VS, Ning Yu, Chen Xing, Can Qin et al.CVPR 2023
- Open-Vocabulary Instance Segmentation via Robust Cross-Modal Pseudo-LabelingDat Huynh, Jason Kuen, Zhe Lin, Jiuxiang Gu et al.CVPR 2022 · 78 citations
- A Simple Framework for Open-Vocabulary Segmentation and DetectionHao Zhang, Feng Li, Xueyan Zou, Shilong Liu et al.ICCV 2023 · 241 citations
- Learning Object-Language Alignments for Open-Vocabulary Object DetectionChuang Lin, Peize Sun, Yi Jiang, Ping Luo et al.ICLR 2023 · 36 citations
- Open-Vocabulary Semantic Segmentation with Mask-adapted CLIPFeng Liang, Bichen Wu, Xiaoliang Dai, Kunpeng Li et al.CVPR 2023
