Learning to Generate Text-Grounded Mask for Open-World Semantic Segmentation from Only Image-Text Pairs
Junbum Cha, Jonghwan Mun, Byungseok Roh
摘要
We tackle open-world semantic segmentation, which aims at learning to segment arbitrary visual concepts in images, by using only image-text pairs without dense annotations. Existing open-world segmentation methods have shown impressive advances by employing contrastive learning (CL) to learn diverse visual concepts and transferring the learned image-level understanding to the segmentation task. However, these CL-based methods suffer from a train-test discrepancy, since it only considers image-text alignment during training, whereas segmentation requires region-text alignment during testing. In this paper, we proposed a novel Text-grounded Contrastive Learning (TCL) framework that enables a model to directly learn region-text alignment. Our method generates a segmentation mask for a given text, extracts text-grounded image embedding from the masked region, and aligns it with text embedding via TCL. By learning region-text alignment directly, our framework encourages a model to directly improve the quality of generated segmentation masks. In addition, for a rigorous and fair comparison, we present a unified evaluation protocol with widely used 8 semantic segmentation datasets. TCL achieves state-of-the-art zero-shot segmentation performances with large margins in all datasets. Code is available at https://github.com/kakaobrain/tcl.
1 This setting is often called both open-world and open-vocabulary. In this paper, we mainly refer to this setting as open-world for clarity.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper59
- Perceptual Grouping in Contrastive Vision-Language ModelsKanchana Ranasinghe, Brandon McKinzie, Sachin Ravi, Yinfei Yang 等ICCV 2023 · 被引用 88 次
- SATR: Zero-Shot Semantic Segmentation of 3D ShapesAhmed Abdelreheem, Ivan Skorokhodov, Maks Ovsjanikov, Peter WonkaICCV 2023 · 被引用 68 次
- Talking to DINO: Bridging Self-Supervised Vision Backbones with Language for Open-Vocabulary SegmentationLuca Barsellotti, Lorenzo Bianchi, Nicola Messina, Fabio Carrara 等ICCV 2025 · 被引用 58 次
- Uncovering Prototypical Knowledge for Weakly Open-Vocabulary Semantic SegmentationFei Zhang, Tianfei Zhou, Boyang Li, Hao He 等NeurIPS 2023 · 被引用 46 次
- EmerDiff: Emerging Pixel-level Semantic Knowledge in Diffusion ModelsKoichi Namekata, Amirmojtaba Sabour, Sanja Fidler, Seung Wook KimICLR 2024 · 被引用 41 次
它引用的顶会 Paper13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 被引用 2,196 次
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 被引用 1,274 次
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve 等ICCV 2021 · 被引用 1,114 次
相关 Paper
- Open Vocabulary Semantic Segmentation with Patch Aligned Contrastive LearningJishnu Mukhoti, Tsung-Yu Lin, Omid Poursaeed, Rui Wang 等CVPR 2023
- Image-Text Co-Decomposition for Text-Supervised Semantic SegmentationJi-Jia Wu, Andy Chia-Hao Chang, Chieh-Yu Chuang, Chun-Pei Chen 等CVPR 2024 · 被引用 6 次
- SegCLIP: Patch Aggregation with Learnable Centers for Open-Vocabulary Semantic SegmentationHuaishao Luo, Junwei Bao, Youzheng Wu, Xiaodong He 等ICML 2023 · 被引用 222 次
- MixReorg: Cross-Modal Mixed Patch Reorganization is a Good Mask Learner for Open-World Semantic SegmentationKaixin Cai, Pengzhen Ren, Yi Zhu, Hang Xu 等ICCV 2023 · 被引用 22 次
- RegionCLIP: Region-based Language-Image PretrainingYiwu Zhong, Jianwei Yang, Pengchuan Zhang, Chunyuan Li 等CVPR 2022 · 被引用 481 次
