Self-Supervised Learning of Object Parts for Semantic Segmentation
Adrian Ziegler, Yuki M. Asano
Abstract
Progress in self-supervised learning has brought strong image representation learning methods. Yet so far, it has mostly focused on image-level learning. In turn, tasks such as unsupervised image segmentation have not benefited from this trend as they require spatially-diverse representations. However, learning dense representations is challenging, as in the unsupervised context it is not clear how to guide the model to learn representations that correspond to various potential object categories. In this paper, we argue that self-supervised learning of object parts is a solution to this issue. Object parts are generalizable: they are a priori independent of an object definition, but can be grouped to form objects a posteriori. To this end, we leverage the recently proposed Vision Transformer's capability of attending to objects and combine it with a spatially dense clustering task for fine-tuning the spatial tokens. Our method surpasses the state-of-the-art on three semantic segmentation benchmarks by 17%-3%, showing that our representations are versatile under various object definitions. Finally, we extend this to fully unsupervised segmentation - which refrains completely from using label information even at test-time - and demonstrate that a simple method for automatically merging discovered object parts based on community detection yields substantial gains..
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c6eb16d1-ee59-4588-8a3c-1b7d7ab1cc66Cited by top-tier papers35
- Self-Supervised Visual Representation Learning with Semantic GroupingXin Wen, Bingchen Zhao, Anlin Zheng, Xiangyu Zhang et al.NeurIPS 2022 · 104 citations
- TagCLIP: A Local-to-Global Framework to Enhance Open-Vocabulary Multi-Label Classification of CLIP without TrainingYuqi Lin, Minghao Chen, Kaipeng Zhang, Hengjia Li et al.AAAI 2024 · 39 citations
- Time Does Tell: Self-Supervised Time-Tuning of Dense Image RepresentationsMohammadreza Salehi, Efstratios Gavves, Cees G. M. Snoek, Yuki M. AsanoICCV 2023 · 34 citations
- MOVE: Unsupervised Movable Object Segmentation and DetectionAdam Bielski, Paolo FavaroNeurIPS 2022 · 30 citations
- MARS: Model-agnostic Biased Object Removal without Additional Supervision for Weakly-Supervised Semantic SegmentationSanghyun Jo, In-Jae Yu, Kyungsu KimICCV 2023 · 29 citations
Builds on22
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
Related papers
- Unsupervised Part Discovery from Contrastive ReconstructionSubhabrata Choudhury, Iro Laina, Christian Rupprecht, Andrea VedaldiNeurIPS 2021 · 74 citations
- Semantic-Aware Superpixel for Weakly Supervised Semantic SegmentationSangtae Kim, Daeyoung Park, Byonghyo ShimAAAI 2023 · 35 citations
- Unsupervised Hierarchical Semantic Segmentation with Multiview Cosegmentation and Clustering TransformersTsung-Wei Ke, Jyh-Jing Hwang, Yunhui Guo, Xudong Wang et al.CVPR 2022 · 34 citations
- FLSL: Feature-level Self-supervised LearningQing Su, Anton Netchaev, Hai Li, Shihao JiNeurIPS 2023 · 9 citations
- GroupViT: Semantic Segmentation Emerges from Text SupervisionJiarui Xu, Shalini De Mello, Sifei Liu, Wonmin Byeon et al.CVPR 2022 · 398 citations
