Segment Anything without Supervision
Xudong Wang, Jingfeng Yang, Trevor Darrell
Abstract
The Segmentation Anything Model (SAM) requires labor-intensive data labeling. We present Unsupervised SAM (UnSAM) for promptable and automatic whole-image segmentation that does not require human annotations. UnSAM utilizes a divide-and-conquer strategy to"discover"the hierarchical structure of visual scenes. We first leverage top-down clustering methods to partition an unlabeled image into instance/semantic level segments. For all pixels within a segment, a bottom-up clustering method is employed to iteratively merge them into larger groups, thereby forming a hierarchical structure. These unsupervised multi-granular masks are then utilized to supervise model training. Evaluated across seven popular datasets, UnSAM achieves competitive results with the supervised counterpart SAM, and surpasses the previous state-of-the-art in unsupervised segmentation by 11% in terms of AR. Moreover, we show that supervised SAM can also benefit from our self-supervised labels. By integrating our unsupervised pseudo masks into SA-1B's ground-truth masks and training UnSAM with only 1% of SA-1B, a lightly semi-supervised UnSAM can often segment entities overlooked by supervised SAM, exceeding SAM's AR by over 6.7% and AP by 3.9% on SA-1B.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 47a56798-974d-430c-90cd-d5221027ce98Cited by top-tier papers13
- TRACE: Your Diffusion Model is Secretly an Instance Edge DetectorSanghyun Jo, Ziseok Lee, Wooyeol Lee, Jonghyun Choi et al.ICLR 2026 · 4 citations
- E-SAM: Training-Free Segment Every Entity ModelWeiming Zhang, Dingwen Xiao, Lei Chen, Lin WangICCV 2025 · 3 citations
- CutS3D: Cutting Semantics in 3D for 2D Unsupervised Instance SegmentationLeon Sick, Dominik Engel, Sebastian Hartwig, Pedro Hermosilla et al.ICCV 2025 · 2 citations
- Scene-Centric Unsupervised Video Panoptic SegmentationChristoph Reich, Oliver Hahn, Nikita Araslanov, Laura Leal-Taixe et al.CVPR 2026 · 1 citation
- S2-UniSeg: Fast Universal Agglomerative Pooling for Scalable Segment Anything Without SupervisionHuihui Xu, Jin Ye, Hongqiu Wang, Changkai Ji et al.AAAI 2026 · 1 citation
Builds on18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
Related papers
- Unleashing the Potential of SAM for Medical Adaptation via Hierarchical DecodingZhiheng Cheng, Qingyue Wei, Hongru Zhu, Yan Wang et al.CVPR 2024
- SOHES: Self-supervised Open-world Hierarchical Entity SegmentationShengcao Cao, Jiuxiang Gu, Jason Kuen, Hao Tan et al.ICLR 2024 · 3 citations
- SAM-CP: Marrying SAM with Composable Prompts for Versatile SegmentationPengfei Chen, Lingxi Xie, Xinyue Huo, Xuehui Yu et al.ICLR 2025
- Segment Anything Model Meets Semi-supervised Medical Image Segmentation: A Novel PerspectiveHaifeng Zhao, Haiyang Li, Lei-Lei Ma, Dengdi SunNeurIPS 2025 · 1 citation
- MaskSAM: Auto-Prompt SAM with Mask Classification for Volumetric Medical Image SegmentationBin Xie, Hao Tang, Bin Duan, Dawen Cai et al.ICCV 2025 · 7 citations
