Unsupervised Object-Level Representation Learning from Scene Images
Jiahao Xie, Xiaohang Zhan, Ziwei Liu, Yew Soon Ong, Chen Change Loy
Abstract
Contrastive self-supervised learning has largely narrowed the gap to supervised pre-training on ImageNet. However, its success highly relies on the object-centric priors of ImageNet, i.e., different augmented views of the same image correspond to the same object. Such a heavily curated constraint becomes immediately infeasible when pre-trained on more complex scene images with many objects. To overcome this limitation, we introduce Object-level Representation Learning (ORL), a new self-supervised learning framework towards scene images. Our key insight is to leverage image-level self-supervised pre-training as the prior to discover object-level semantic correspondence, thus realizing object-level representation learning from scene images. Extensive experiments on COCO show that ORL significantly improves the performance of self-supervised learning on scene images, even surpassing supervised ImageNet pre-training on several downstream tasks. Furthermore, ORL improves the downstream performance when more unlabeled scene images are available, demonstrating its great potential of harnessing unlabeled data in the wild. We hope our approach can motivate future research on more general-purpose unsupervised representation learning from scene data. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dda22b8f-dea3-4ce9-8e78-ddd9681da22cCited by top-tier papers5
- Point-Level Region Contrast for Object Detection Pre-TrainingYutong Bai, Xinlei Chen, Alexander Kirillov, Alan L. Yuille et al.CVPR 2022 · 43 citations
- Multi-Label Self-Supervised Learning with Scene ImagesKe Zhu, Minghao Fu, Jianxin WuICCV 2023 · 21 citations
- Scene Consistency Representation Learning for Video Scene SegmentationHaoqian Wu, Keyu Chen, Yanan Luo, Ruizhi Qiao et al.CVPR 2022 · 19 citations
- R-MAE: Regions Meet Masked AutoencodersDuy-Kien Nguyen, Yanghao Li, Vaibhav Aggarwal, Martin R. Oswald et al.ICLR 2024 · 18 citations
- FORLA: Federated Object-Centric Representation Learning with Slot AttentionGuiqiu Liao, Matjaz Jogan, Eric Eaton, Daniel A. HashimotoNeurIPS 2025 · 3 citations
Builds on23
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- What Makes for Good Views for Contrastive Learning?Yonglong Tian, Chen Sun, Ben Poole, Dilip Krishnan et al.NeurIPS 2020 · 1,631 citations
- Data-Efficient Image Recognition with Contrastive Predictive CodingOlivier J. HénaffICML 2020 · 1,553 citations
Related papers
- UniVIP: A Unified Framework for Self-Supervised Visual Pre-trainingZhaowen Li, Yousong Zhu, Fan Yang, Wei Li et al.CVPR 2022 · 29 citations
- Demystifying Contrastive Self-Supervised Learning: Invariances, Augmentations and Dataset BiasesSenthil Purushwalkam, Abhinav GuptaNeurIPS 2020 · 240 citations
- CASTing Your Model: Learning To Localize Improves Self-Supervised RepresentationsRamprasaath R. Selvaraju, Karan Desai, Justin Johnson, Nikhil NaikCVPR 2021
- Hyperbolic Contrastive Learning for Visual Representations beyond ObjectsSongwei Ge, Shlok Mishra, Simon Kornblith, Chun-Liang Li et al.CVPR 2023
- Self-Supervised Visual Representation Learning with Semantic GroupingXin Wen, Bingchen Zhao, Anlin Zheng, Xiangyu Zhang et al.NeurIPS 2022 · 104 citations
