SegGPT: Towards Segmenting Everything In Context
Xinlong Wang, Xiaosong Zhang, Yue Cao, Wen Wang, Chunhua Shen, Tiejun Huang
Abstract
We present SegGPT, a generalist model for segmenting everything in context. We unify various segmentation tasks into a generalist in-context learning framework that accommodates different kinds of segmentation data by transforming them into the same format of images. The training of SegGPT is formulated as an in-context coloring problem with random color mapping for each data sample. The objective is to accomplish diverse tasks according to the context, rather than relying on specific colors. After training, SegGPT can perform arbitrary segmentation tasks in images or videos via in-context inference, such as object instance, stuff, part, contour, and text. SegGPT is evaluated on a broad range of tasks, including few-shot semantic segmentation, video object segmentation, semantic segmentation, and panoptic segmentation. Our results show strong capabilities in segmenting in-domain and out-of-domain targets, either qualitatively or quantitatively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9b013227-2acc-4bbc-b93f-877e8ba87980Cited by top-tier papers29
- A Survey on In-context LearningQingxiu Dong, Lei Li, Damai Dai, Ce Zheng et al.EMNLP 2024 · 479 citations
- Bridge the Points: Graph-based Few-shot Segment Anything SemanticallyAnqi Zhang, Guangyu Gao, Jianbo Jiao, Chi Harold Liu et al.NeurIPS 2024 · 56 citations
- VRP-SAM: SAM with Visual Reference PromptYanpeng Sun, Jiahui Chen, Shan Zhang, Xinyu Zhang et al.CVPR 2024 · 49 citations
- Spider: A Unified Framework for Context-dependent Concept SegmentationXiaoqi Zhao, Youwei Pang, Wei Ji, Baicheng Sheng et al.ICML 2024 · 21 citations
- AWRaCLe: All-Weather Image Restoration Using Visual In-Context LearningSudarshan Rajagopalan, Vishal M. PatelAAAI 2025 · 20 citations
Builds on23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- YOLACT: Real-Time Instance SegmentationDaniel Bolya, Chong Zhou, Fanyi Xiao, Yong Jae LeeICCV 2019 · 2,075 citations
- SOLOv2: Dynamic and Fast Instance SegmentationXinlong Wang, Rufeng Zhang, Tao Kong, Lei Li et al.NeurIPS 2020 · 1,193 citations
Related papers
- OMG-Seg: Is One Model Good Enough for all Segmentation?Xiangtai Li, Haobo Yuan, Wei Li, Henghui Ding et al.CVPR 2024
- FreeSeg: Unified, Universal and Open-Vocabulary Image SegmentationJie Qin, Jie Wu, Pengxiang Yan, Ming Li et al.CVPR 2023
- TarViS: A Unified Approach for Target-Based Video SegmentationAli Athar, Alexander Hermans, Jonathon Luiten, Deva Ramanan et al.CVPR 2023
- MasQCLIP for Open-Vocabulary Universal Image SegmentationXin Xu, Tianyi Xiong, Zheng Ding, Zhuowen TuICCV 2023 · 57 citations
- Unsupervised Universal Image SegmentationDantong Niu, Xudong Wang, Xinyang Han, Long Lian et al.CVPR 2024 · 29 citations
