HEAP: Unsupervised Object Discovery and Localization with Contrastive Grouping
Xin Zhang, Jinheng Xie, Yuan Yuan, Michael Bi Mi, Robby T. Tan
摘要
Unsupervised object discovery and localization aims to detect or segment objects in an image without any supervision. Recent efforts have demonstrated a notable potential to identify salient foreground objects by utilizing self-supervised transformer features. However, their scopes only build upon patch-level features within an image, neglecting region/image-level and cross-image relationships at a broader scale. Moreover, these methods cannot differentiate various semantics from multiple instances. To address these problems, we introduce Hierarchical mErging framework via contrAstive grouPing (HEAP). Specifically, a novel lightweight head with cross-attention mechanism is designed to adaptively group intra-image patches into semantically coherent regions based on correlation among self-supervised features. Further, to ensure the distinguishability among various regions, we introduce a region-level contrastive clustering loss to pull closer similar regions across images. Also, an image-level contrastive loss is present to push foreground and background representations apart, with which foreground objects and background are accordingly discovered. HEAP facilitates efficient hierarchical image decomposition, which contributes to more accurate object discovery while also enabling differentiation among objects of various classes. Extensive experimental results on semantic segmentation retrieval, unsupervised object discovery, and saliency detection tasks demonstrate that HEAP achieves state-of-the-art performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- 3DOT: Texture Transfer for 3DGS Objects from a Single Reference ImageXiao Cao, Beibei Lin, Bo Wang, Zhiyong Huang 等NeurIPS 2025 · 被引用 8 次
- Object Concepts Emerge from MotionHaoqian Liang, Xiaohui Wang, Zhichao Li, Ya Yang 等NeurIPS 2025
- Synthetic-to-Real Self-supervised Robust Depth Estimation via Learning with Motion and Structure PriorsWeilong Yan, Ming Li, Haipeng Li, Shuwei Shao 等CVPR 2025
- unMORE: Unsupervised Multi-Object Segmentation via Center-Boundary ReasoningYafei Yang, Zihui Zhang, Bo YangICML 2025
它引用的顶会 Paper23
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 被引用 2,340 次
- Segmenter: Transformer for Semantic SegmentationRobin Strudel, Ricardo Garcia, Ivan Laptev, Cordelia SchmidICCV 2021 · 被引用 1,898 次
相关 Paper
- Unsupervised Salient Instance DetectionXin Tian, Ke Xu, Rynson W. H. LauCVPR 2024 · 被引用 2 次
- Unsupervised Hierarchical Semantic Segmentation with Multiview Cosegmentation and Clustering TransformersTsung-Wei Ke, Jyh-Jing Hwang, Yunhui Guo, Xudong Wang 等CVPR 2022 · 被引用 34 次
- Self-Supervised Transformers for Unsupervised Object Discovery using Normalized CutYangtao Wang, Xi Shen, Shell Xu Hu, Yuan Yuan 等CVPR 2022 · 被引用 143 次
- Self-Supervised Visual Representation Learning from Hierarchical GroupingXiao Zhang, Michael MaireNeurIPS 2020 · 被引用 82 次
- Self-Supervised Learning of Object Parts for Semantic SegmentationAdrian Ziegler, Yuki M. AsanoCVPR 2022 · 被引用 87 次
