CuVLER: Enhanced Unsupervised Object Discoveries through Exhaustive Self-Supervised Transformers
Shahaf Arica, Or Rubin, Sapir Gershov, Shlomi Laufer
Abstract
In this paper, we introduce VoteCut, an innovative method for unsupervised object discovery that leverages feature representations from multiple self-supervised models. VoteCut employs normalized-cut based graph partitioning, clustering and a pixel voting approach. Additionally, We present CuVLER (Cut-Vote-and-LEaRn), a zero-shot model, trained using pseudo-labels, generated by VoteCut, and a novel soft target loss to refine segmentation accuracy. Through rigorous evaluations across multiple datasets and several unsupervised setups, our methods demonstrate significant improvements in comparison to previous state-ofthe-art models. Our ablation studies further highlight the contributions of each component, revealing the robustness and efficacy of our approach. Collectively, VoteCut and CuVLER pave the way for future advancements in image segmentation. The project code is available on GitHub at https://github.com/shahaf-arica/CuVLER
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- Hierarchy-Agnostic Unsupervised Segmentation: Parsing Semantic Image StructureSimone Rossetti, Fiora PirriNeurIPS 2024 · 2 citations
- CutS3D: Cutting Semantics in 3D for 2D Unsupervised Instance SegmentationLeon Sick, Dominik Engel, Sebastian Hartwig, Pedro Hermosilla et al.ICCV 2025 · 2 citations
- Beyond Single Images: Retrieval Self-Augmented Unsupervised Camouflaged Object DetectionJi Du, Xin Wang, Fangwei Hao, Mingyang Yu et al.ICCV 2025 · 2 citations
- Scene-Centric Unsupervised Video Panoptic SegmentationChristoph Reich, Oliver Hahn, Nikita Araslanov, Laura Leal-Taixe et al.CVPR 2026 · 1 citation
- S2-UniSeg: Fast Universal Agglomerative Pooling for Scalable Segment Anything Without SupervisionHuihui Xu, Jin Ye, Hongqiu Wang, Changkai Ji et al.AAAI 2026 · 1 citation
Builds on13
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
Related papers
- Self-Supervised Transformers for Unsupervised Object Discovery using Normalized CutYangtao Wang, Xi Shen, Shell Xu Hu, Yuan Yuan et al.CVPR 2022 · 143 citations
- Unsupervised Universal Image SegmentationDantong Niu, Xudong Wang, Xinyang Han, Long Lian et al.CVPR 2024 · 29 citations
- Ensemble Foreground Management for Unsupervised Object DiscoveryZiling Wu, Armaghan Moemeni, Praminda Caleb-SollyICCV 2025 · 1 citation
- VideoCutLER: Surprisingly Simple Unsupervised Video Instance SegmentationXudong Wang, Ishan Misra, Ziyun Zeng, Rohit Girdhar et al.CVPR 2024
- Unsupervised Semantic Segmentation with Self-supervised Object-centric RepresentationsAndrii Zadaianchuk, Matthäus Kleindessner, Yi Zhu, Francesco Locatello et al.ICLR 2023 · 16 citations
