Ensemble Foreground Management for Unsupervised Object Discovery
Ziling Wu, Armaghan Moemeni, Praminda Caleb-Solly
Abstract
Unsupervised object discovery (UOD) aims to detect and segment objects in 2D images without handcrafted anno- tations. Recent progress in self-supervised representation learning [9, 66] has led to some success in UOD algo- rithms [21, 53, 69]. However, the absence of ground truth provides existing UOD methods with two challenges: 1) determining if a discovered region is foreground or back- ground, and 2) knowing how many objects remain undiscov- ered. To address these two problems, previous solutions rely on foreground priors [53, 59, 67, 69] to distinguish if the discovered region is foreground, and conduct one or fixed it- erations of discovery. However, the existing foreground pri- ors are heuristic and not always robust, and a fixed number of discoveries leads to under or over-segmentation, since the number of objects in images varies. This paper intro- duces UnionCut, a robust and well-grounded foreground prior based on min-cut [2] and ensemble methods [18] that detects the union of foreground areas of an image, allow- ing UOD algorithms to identify foreground objects and stop discovery once the majority of the foreground union in the image is segmented. In addition, we propose UnionSeg, a distilled transformer of UnionCut that outputs the fore- ground union more efficiently and accurately. Our experiments show that by combining with UnionCut or UnionSeg, previous state-of-the-art UOD methods [21, 53, 54, 69] wit- ness an increase in the performance of single object discovery, saliency detection and self-supervised instance seg- mentation on various benchmarks. The code is available at https://github.com/YFaris/UnionCut.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7dbf8a4f-be46-444f-8426-c036efb78a98Builds on26
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
Related papers
- Self-Supervised Transformers for Unsupervised Object Discovery using Normalized CutYangtao Wang, Xi Shen, Shell Xu Hu, Yuan Yuan et al.CVPR 2022 · 143 citations
- MOVE: Unsupervised Movable Object Segmentation and DetectionAdam Bielski, Paolo FavaroNeurIPS 2022 · 30 citations
- CuVLER: Enhanced Unsupervised Object Discoveries through Exhaustive Self-Supervised TransformersShahaf Arica, Or Rubin, Sapir Gershov, Shlomi LauferCVPR 2024
- UNION: Unsupervised 3D Object Detection using Object Appearance-based Pseudo-ClassesTed de Vries Lentsch, Holger Caesar, Dariu GavrilaNeurIPS 2024 · 30 citations
- Unsupervised Semantic Segmentation with Self-supervised Object-centric RepresentationsAndrii Zadaianchuk, Matthäus Kleindessner, Yi Zhu, Francesco Locatello et al.ICLR 2023 · 16 citations
