Cut and Learn for Unsupervised Object Detection and Instance Segmentation
Xudong Wang, Rohit Girdhar, Stella X. Yu, Ishan Misra
摘要
OpenImages Datasets AP 50 Figure 1 . Zero-shot unsupervised object detection and instance segmentation using our CutLER model, which is trained without human supervision. We evaluate the model using the standard detection AP box 50 . CutLER gives a strong performance on a variety of benchmarks spanning diverse image domains -video frames, paintings, clip arts, complex scenes, etc. Compared to the previous stateof-the-art method, FreeSOLO [47] with a backbone of ResNet101, CutLER with a backbone of ResNet50 provides strong gains on all benchmarks, increasing performance by more than 2× on 10 of the 11 benchmarks. We evaluate [47] with its official code and checkpoint.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper78
- Vision Transformers Need RegistersTimothée Darcet, Maxime Oquab, Julien Mairal, Piotr BojanowskiICLR 2024 · 被引用 769 次
- Object-Centric Slot DiffusionJindong Jiang, Fei Deng, Gautam Singh, Sungjin AhnNeurIPS 2023 · 被引用 106 次
- A Touch, Vision, and Language Dataset for Multimodal AlignmentLetian Fu, Gaurav Datta, Huang Huang, William Chung-Ho Panitch 等ICML 2024 · 被引用 89 次
- Self-supervised Object-Centric Learning for VideosGörkay Aydemir, Weidi Xie, Fatma GüneyNeurIPS 2023 · 被引用 61 次
- Talking to DINO: Bridging Self-Supervised Vision Backbones with Language for Open-Vocabulary SegmentationLuca Barsellotti, Lorenzo Bianchi, Nicola Messina, Fabio Carrara 等ICCV 2025 · 被引用 58 次
它引用的顶会 Paper21
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
相关 Paper
- VideoCutLER: Surprisingly Simple Unsupervised Video Instance SegmentationXudong Wang, Ishan Misra, Ziyun Zeng, Rohit Girdhar 等CVPR 2024
- Unsupervised Universal Image SegmentationDantong Niu, Xudong Wang, Xinyang Han, Long Lian 等CVPR 2024 · 被引用 29 次
- FreeSOLO: Learning to Segment Objects without AnnotationsXinlong Wang, Zhiding Yu, Shalini De Mello, Jan Kautz 等CVPR 2022 · 被引用 100 次
- CuVLER: Enhanced Unsupervised Object Discoveries through Exhaustive Self-Supervised TransformersShahaf Arica, Or Rubin, Sapir Gershov, Shlomi LauferCVPR 2024
- ZBS: Zero-Shot Background Subtraction via Instance-Level Background Modeling and Foreground SelectionYongqi An, Xu Zhao, Tao Yu, Haiyun Gu 等CVPR 2023
