Unsupervised Semantic Segmentation by Distilling Feature Correspondences
Mark Hamilton, Zhoutong Zhang, Bharath Hariharan, Noah Snavely, William T. Freeman
Abstract
Unsupervised semantic segmentation aims to discover and localize semantically meaningful categories within image corpora without any form of annotation. To solve this task, algorithms must produce features for every pixel that are both semantically meaningful and compact enough to form distinct clusters. Unlike previous works which achieve this with a single end-to-end framework, we propose to separate feature learning from cluster compactification. Empirically, we show that current unsupervised feature learning frameworks already generate dense features whose correlations are semantically consistent. This observation motivates us to design STEGO (elf-supervised ransformer with nergy-based raph ptimization), a novel framework that distills unsupervised features into high-quality discrete semantic labels. At the core of STEGO is a novel contrastive loss function that encourages features to form compact clusters while preserving their relationships across the corpora. STEGO yields a significant improvement over the prior state of the art, on both the CocoStuff () and Cityscapes () semantic segmentation challenges.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers98
- Emergent Correspondence from Image DiffusionLuming Tang, Menglin Jia, Qianqian Wang, Cheng Perng Phoo et al.NeurIPS 2023 · 555 citations
- Diffusion Hyperfeatures: Searching Through Time and Space for Semantic CorrespondenceGrace Luo, Lisa Dunlap, Dong Huk Park, Aleksander Holynski et al.NeurIPS 2023 · 261 citations
- Segment Any 3D GaussiansJiazhong Cen, Jiemin Fang, Chen Yang, Lingxi Xie et al.AAAI 2025 · 175 citations
- Weakly Supervised 3D Open-vocabulary SegmentationKunhao Liu, Fangneng Zhan, Jiahui Zhang, Muyu Xu et al.NeurIPS 2023 · 173 citations
- ReCo: Retrieve and Co-segment for Zero-shot TransferGyungin Shin, Weidi Xie, Samuel AlbanieNeurIPS 2022 · 160 citations
Builds on11
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- Big Self-Supervised Models are Strong Semi-Supervised LearnersTing Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi et al.NeurIPS 2020 · 2,611 citations
Related papers
- SmooSeg: Smoothness Prior for Unsupervised Semantic SegmentationMengcheng Lan, Xinjiang Wang, Yiping Ke, Jiaxing Xu et al.NeurIPS 2023 · 28 citations
- Unsupervised Universal Image SegmentationDantong Niu, Xudong Wang, Xinyang Han, Long Lian et al.CVPR 2024 · 29 citations
- EAGLE: Eigen Aggregation Learning for Object-Centric Unsupervised Semantic SegmentationChanyoung Kim, Woojung Han, Dayun Ju, Seong Jae HwangCVPR 2024
- Integrating Low-Level Visual Cues for Enhanced Unsupervised Semantic SegmentationYuhao Qing, Dan Zeng, Shaorong Xie, Kaer Huang et al.AAAI 2025 · 1 citation
- Unsupervised Semantic Segmentation with Self-supervised Object-centric RepresentationsAndrii Zadaianchuk, Matthäus Kleindessner, Yi Zhu, Francesco Locatello et al.ICLR 2023 · 16 citations
