Language-driven Semantic Segmentation
Boyi Li, Kilian Q. Weinberger, Serge J. Belongie, Vladlen Koltun, René Ranftl
Abstract
We present LSeg, a novel model for language-driven semantic image segmentation. LSeg uses a text encoder to compute embeddings of descriptive input labels (e.g., "grass" or "building") together with a transformer-based image encoder that computes dense per-pixel embeddings of the input image. The image encoder is trained with a contrastive objective to align pixel embeddings to the text embedding of the corresponding semantic class. The text embeddings provide a flexible label representation in which semantically similar labels map to similar regions in the embedding space (e.g., "cat" and "furry"). This allows LSeg to generalize to previously unseen categories at test time, without retraining or even requiring a single additional training sample. We demonstrate that our approach achieves highly competitive zero-shot performance compared to existing zero-and few-shot semantic segmentation methods, and even matches the accuracy of traditional segmentation algorithms when a fixed label set is provided. Code and demo are available at https://github.com/isl-org/lang-seg .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers328
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- LERF: Language Embedded Radiance FieldsJustin Kerr, Chung Min Kim, Ken Goldberg, Angjoo Kanazawa et al.ICCV 2023 · 620 citations
- Instruct-NeRF2NeRF: Editing 3D Scenes with InstructionsAyaan Haque, Matthew Tancik, Alexei A. Efros, Aleksander Holynski et al.ICCV 2023 · 544 citations
- Decomposing NeRF for Editing via Feature Field DistillationSosuke Kobayashi, Eiichi Matsumoto, Vincent SitzmannNeurIPS 2022 · 479 citations
- Image Segmentation Using Text and Image PromptsTimo Lüddecke, Alexander S. EckerCVPR 2022 · 457 citations
Builds on21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- PANet: Few-Shot Image Semantic Segmentation With Prototype AlignmentKaixin Wang, Jun Hao Liew, Yingtian Zou, Daquan Zhou et al.ICCV 2019 · 1,404 citations
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve et al.ICCV 2021 · 1,114 citations
Related papers
- ZeroCap: Zero-Shot Image-to-Text Generation for Visual-Semantic ArithmeticYoad Tewel, Yoav Shalev, Idan Schwartz, Lior WolfCVPR 2022 · 129 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- ReSTR: Convolution-free Referring Image Segmentation Using TransformersNamyup Kim, Dongwon Kim, Suha Kwak, Cuiling Lan et al.CVPR 2022 · 149 citations
- Context-aware Feature Generation For Zero-shot Semantic SegmentationZhangxuan Gu, Siyuan Zhou, Li Niu, Zihan Zhao et al.ACM MM 2020 · 111 citations
- Cap2Seg: Inferring Semantic and Spatial Context from Captions for Zero-Shot Image SegmentationGuiyu Tian, Shuai Wang, Jie Feng, Li Zhou et al.ACM MM 2020 · 13 citations
