Learning Open-Vocabulary Semantic Segmentation Models From Natural Language Supervision
Jilan Xu, Junlin Hou, Yuejie Zhang, Rui Feng, Yi Wang, Yu Qiao, Weidi Xie
摘要
In this paper, we consider the problem of openvocabulary semantic segmentation (OVS), which aims to segment objects of arbitrary classes instead of pre-defined, closed-set categories. The main contributions are as follows: First, we propose a transformer-based model for OVS, termed as OVSegmentor, which only exploits webcrawled image-text pairs for pre-training without using any mask annotations. OVSegmentor assembles the image pixels into a set of learnable group tokens via a slot-attention based binding module, and aligns the group tokens to the corresponding caption embedding. Second, we propose two proxy tasks for training, namely masked entity completion and cross-image mask consistency. The former aims to infer all masked entities in the caption given the group tokens, that enables the model to learn fine-grained alignment between visual groups and text entities. The latter enforces consistent mask predictions between images that contain shared entities, which encourages the model to learn visual invariance. Third, we construct CC4M dataset for pre-training by filtering CC12M with frequently appeared entities, which significantly improves training efficiency. Fourth, we perform zero-shot transfer on three benchmark datasets, PASCAL VOC 2012, PASCAL Context, and COCO Object. Our model achieves superior segmentation results over the state-of-the-art method by using only 3% data (4M vs 134M) for pre-training. Code and pre-trained models will be released for future research.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper57
- Perceptual Grouping in Contrastive Vision-Language ModelsKanchana Ranasinghe, Brandon McKinzie, Sachin Ravi, Yinfei Yang 等ICCV 2023 · 被引用 88 次
- Hierarchical Open-vocabulary Universal Image SegmentationXudong Wang, Shufan Li, Konstantinos Kallidromitis, Yusuke Kato 等NeurIPS 2023 · 被引用 74 次
- SATR: Zero-Shot Semantic Segmentation of 3D ShapesAhmed Abdelreheem, Ivan Skorokhodov, Maks Ovsjanikov, Peter WonkaICCV 2023 · 被引用 68 次
- Talking to DINO: Bridging Self-Supervised Vision Backbones with Language for Open-Vocabulary SegmentationLuca Barsellotti, Lorenzo Bianchi, Nicola Messina, Fabio Carrara 等ICCV 2025 · 被引用 58 次
- Improving fine-grained understanding in image-text pre-trainingIoana Bica, Anastasija Ilic, Matthias Bauer, Goker Erdogan 等ICML 2024 · 被引用 53 次
它引用的顶会 Paper32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty 等NeurIPS 2021 · 被引用 2,985 次
相关 Paper
- Open-Vocabulary Universal Image Segmentation with MaskCLIPZheng Ding, Jieke Wang, Zhuowen TuICML 2023 · 被引用 150 次
- A Simple Framework for Open-Vocabulary Segmentation and DetectionHao Zhang, Feng Li, Xueyan Zou, Shilong Liu 等ICCV 2023 · 被引用 241 次
- MixReorg: Cross-Modal Mixed Patch Reorganization is a Good Mask Learner for Open-World Semantic SegmentationKaixin Cai, Pengzhen Ren, Yi Zhu, Hang Xu 等ICCV 2023 · 被引用 22 次
- SegCLIP: Patch Aggregation with Learnable Centers for Open-Vocabulary Semantic SegmentationHuaishao Luo, Junwei Bao, Youzheng Wu, Xiaodong He 等ICML 2023 · 被引用 222 次
- Emergent Open-Vocabulary Semantic Segmentation from Off-the-Shelf Vision-Language ModelsJiayun Luo, Siddhesh Khandelwal, Leonid Sigal, Boyang LiCVPR 2024 · 被引用 10 次
