Cut and Learn for Unsupervised Object Detection and Instance Segmentation
Xudong Wang, Rohit Girdhar, Stella X. Yu, Ishan Misra
Abstract
OpenImages Datasets AP 50 Figure 1 . Zero-shot unsupervised object detection and instance segmentation using our CutLER model, which is trained without human supervision. We evaluate the model using the standard detection AP box 50 . CutLER gives a strong performance on a variety of benchmarks spanning diverse image domains -video frames, paintings, clip arts, complex scenes, etc. Compared to the previous stateof-the-art method, FreeSOLO [47] with a backbone of ResNet101, CutLER with a backbone of ResNet50 provides strong gains on all benchmarks, increasing performance by more than 2× on 10 of the 11 benchmarks. We evaluate [47] with its official code and checkpoint.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 237658be-c40f-4e00-9422-1e69a32bf9f9Cited by top-tier papers78
- Vision Transformers Need RegistersTimothée Darcet, Maxime Oquab, Julien Mairal, Piotr BojanowskiICLR 2024 · 769 citations
- Object-Centric Slot DiffusionJindong Jiang, Fei Deng, Gautam Singh, Sungjin AhnNeurIPS 2023 · 106 citations
- A Touch, Vision, and Language Dataset for Multimodal AlignmentLetian Fu, Gaurav Datta, Huang Huang, William Chung-Ho Panitch et al.ICML 2024 · 89 citations
- Self-supervised Object-Centric Learning for VideosGörkay Aydemir, Weidi Xie, Fatma GüneyNeurIPS 2023 · 61 citations
- Talking to DINO: Bridging Self-Supervised Vision Backbones with Language for Open-Vocabulary SegmentationLuca Barsellotti, Lorenzo Bianchi, Nicola Messina, Fabio Carrara et al.ICCV 2025 · 58 citations
Builds on21
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
Related papers
- VideoCutLER: Surprisingly Simple Unsupervised Video Instance SegmentationXudong Wang, Ishan Misra, Ziyun Zeng, Rohit Girdhar et al.CVPR 2024
- Unsupervised Universal Image SegmentationDantong Niu, Xudong Wang, Xinyang Han, Long Lian et al.CVPR 2024 · 29 citations
- FreeSOLO: Learning to Segment Objects without AnnotationsXinlong Wang, Zhiding Yu, Shalini De Mello, Jan Kautz et al.CVPR 2022 · 100 citations
- CuVLER: Enhanced Unsupervised Object Discoveries through Exhaustive Self-Supervised TransformersShahaf Arica, Or Rubin, Sapir Gershov, Shlomi LauferCVPR 2024
- ZBS: Zero-Shot Background Subtraction via Instance-Level Background Modeling and Foreground SelectionYongqi An, Xu Zhao, Tao Yu, Haiyun Gu et al.CVPR 2023
