Shatter and Gather: Learning Referring Image Segmentation with Text Supervision
Dongwon Kim, Namyup Kim, Cuiling Lan, Suha Kwak
Abstract
Referring image segmentation, the task of segmenting any arbitrary entities described in free-form texts, opens up a variety of vision applications. However, manual labeling of training data for this task is prohibitively costly, leading to lack of labeled data for training. We address this issue by a weakly supervised learning approach using text descriptions of training images as the only source of supervision. To this end, we first present a new model that discovers semantic entities in input image and then combines such entities relevant to text query to predict the mask of the referent. We also present a new loss function that allows the model to be trained without any further supervision. Our method was evaluated on four public benchmarks for referring image segmentation, where it clearly outperformed the existing method for the same task and recent open-vocabulary segmentation models on all the benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 00a9fb84-5674-4799-b9f6-ee34f8e99f07Cited by top-tier papers6
- SAM as the Guide: Mastering Pseudo-Label Refinement in Semi-Supervised Referring Expression SegmentationDanni Yang, Jiayi Ji, Yiwei Ma, Tianyu Guo et al.ICML 2024 · 19 citations
- Boosting Weakly Supervised Referring Image Segmentation via Progressive ComprehensionZaiquan Yang, Yuhao Liu, Jiaying Lin, Gerhard P. Hancke et al.NeurIPS 2024 · 14 citations
- Bootstrapping Top-down Information for Self-modulating Slot AttentionDongwon Kim, Seoyeon Kim, Suha KwakNeurIPS 2024 · 7 citations
- CTRL-O: Language-Controllable Object-Centric Visual Representation LearningAniket Didolkar, Andrii Zadaianchuk, Rabiul Awal, Maximilian Seitzer et al.CVPR 2025
- Curriculum Point Prompting for Weakly-Supervised Referring Image SegmentationQiyuan Dai, Sibei YangCVPR 2024
Builds on34
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran et al.NeurIPS 2020 · 1,275 citations
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve et al.ICCV 2021 · 1,114 citations
Related papers
- CLIP as RNN: Segment Countless Visual Concepts without Training EndeavorShuyang Sun, Runjia Li, Philip Torr, Xiuye Gu et al.CVPR 2024 · 22 citations
- ReSTR: Convolution-free Referring Image Segmentation Using TransformersNamyup Kim, Dongwon Kim, Suha Kwak, Cuiling Lan et al.CVPR 2022 · 149 citations
- Referring Image Segmentation Using Text SupervisionFang Liu, Yuhao Liu, Yuqiu Kong, Ke Xu et al.ICCV 2023 · 52 citations
- Zero-shot Referring Image Segmentation with Global-Local Context FeaturesSeonghoon Yu, Paul Hongsuck Seo, Jeany SonCVPR 2023
- Learning Open-Vocabulary Semantic Segmentation Models From Natural Language SupervisionJilan Xu, Junlin Hou, Yuejie Zhang, Rui Feng et al.CVPR 2023
