Self-Supervised Visual Representation Learning from Hierarchical Grouping
Xiao Zhang, Michael Maire
Abstract
We create a framework for bootstrapping visual representation learning from a primitive visual grouping capability. We operationalize grouping via a contour detector that partitions an image into regions, followed by merging of those regions into a tree hierarchy. A small supervised dataset suffices for training this grouping primitive. Across a large unlabeled dataset, we apply this learned primitive to automatically predict hierarchical region structure. These predictions serve as guidance for self-supervised contrastive feature learning: we task a deep network with producing per-pixel embeddings whose pairwise distances respect the region hierarchy. Experiments demonstrate that our approach can serve as state-of-the-art generic pre-training, benefiting downstream tasks. We additionally explore applications to semantic region search and video-based object instance tracking.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b06d1f00-2e00-468d-8957-baccf93c73fbCited by top-tier papers35
- Unsupervised Semantic Segmentation by Contrasting Object Mask ProposalsWouter Van Gansbeke, Simon Vandenhende, Stamatios Georgoulis, Luc Van GoolICCV 2021 · 285 citations
- MST: Masked Self-Supervised Transformer for Visual RepresentationZhaowen Li, Zhiyang Chen, Fan Yang, Wei Li et al.NeurIPS 2021 · 194 citations
- Efficient Visual Pretraining with Contrastive DetectionOlivier J. Hénaff, Skanda Koppula, Jean-Baptiste Alayrac, Aäron van den Oord et al.ICCV 2021 · 186 citations
- ReCo: Retrieve and Co-segment for Zero-shot TransferGyungin Shin, Weidi Xie, Samuel AlbanieNeurIPS 2022 · 160 citations
- Region-aware Contrastive Learning for Semantic SegmentationHanzhe Hu, Jinshi Cui, Liwei WangICCV 2021 · 132 citations
Builds on4
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Local Aggregation for Unsupervised Learning of Visual EmbeddingsChengxu Zhuang, Alex Lin Zhai, Daniel YaminsICCV 2019 · 462 citations
- SegSort: Segmentation by Discriminative Sorting of SegmentsJyh-Jing Hwang, Stella X. Yu, Jianbo Shi, Maxwell D. Collins et al.ICCV 2019 · 160 citations
- Momentum Contrast for Unsupervised Visual Representation LearningKaiming He, Haoqi Fan, Yuxin Wu, Saining Xie et al.CVPR 2020
Related papers
- Self-Supervised Visual Representation Learning with Semantic GroupingXin Wen, Bingchen Zhao, Anlin Zheng, Xiangyu Zhang et al.NeurIPS 2022 · 104 citations
- Dense Semantic Contrast for Self-Supervised Visual Representation LearningXiaoni Li, Yu Zhou, Yifei Zhang, Aoting Zhang et al.ACM MM 2021 · 35 citations
- Point-Level Region Contrast for Object Detection Pre-TrainingYutong Bai, Xinlei Chen, Alexander Kirillov, Alan L. Yuille et al.CVPR 2022 · 43 citations
- Train a One-Million-Way Instance Classifier for Unsupervised Visual Representation LearningYu Liu, Lianghua Huang, Pan Pan, Bin Wang et al.AAAI 2021 · 3 citations
- Object Concepts Emerge from MotionHaoqian Liang, Xiaohui Wang, Zhichao Li, Ya Yang et al.NeurIPS 2025
