MOST: Multiple Object localization with Self-supervised Transformers for object discovery
Sai Saketh Rambhatla, Ishan Misra, Rama Chellappa, Abhinav Shrivastava
摘要
We tackle the challenging task of unsupervised object localization in this work. Recently, transformers trained with self-supervised learning have been shown to exhibit object localization properties without being trained for this task. In this work, we present Multiple Object localization with Self-supervised Transformers (MOST) that uses features of transformers trained using self-supervised learning to localize multiple objects in real world images. MOST analyzes the similarity maps of the features using box counting; a fractal analysis tool to identify tokens lying on foreground patches. The identified tokens are then clustered together, and tokens of each cluster are used to generate bounding boxes on foreground regions. Unlike recent state-of-the-art object localization methods, MOST can localize multiple objects per image and outperforms SOTA algorithms on several object localization and discovery benchmarks on PASCAL-VOC 07, 12 and COCO20k datasets. Additionally, we show that MOST can be used for self-supervised pretraining of object detectors, and yields consistent improvements on fully, semi-supervised object detection and unsupervised region proposal generation.Our project is publicly available at rssaketh.github.io/most.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- DiPEx: Dispersing Prompt Expansion for Class-Agnostic Object DetectionJia Syuen Lim, Zhuoxiao Chen, Zhi Chen, Mahsa Baktashmotlagh 等NeurIPS 2024 · 被引用 19 次
- Ensemble Foreground Management for Unsupervised Object DiscoveryZiling Wu, Armaghan Moemeni, Praminda Caleb-SollyICCV 2025 · 被引用 1 次
- Trokens: Semantic-Aware Relational Trajectory Tokens for Few-Shot Action RecognitionPulkit Kumar, Shuaiyi Huang, Matthew Walmer, Sai Saketh Rambhatla 等ICCV 2025
它引用的顶会 Paper15
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- With a Little Help from My Friends: Nearest-Neighbor Contrastive Learning of Visual RepresentationsDebidatta Dwibedi, Yusuf Aytar, Jonathan Tompson, Pierre Sermanet 等ICCV 2021 · 被引用 542 次
- Learning to Discover Novel Visual Categories via Deep Transfer ClusteringKai Han, Andrea Vedaldi, Andrew ZissermanICCV 2019 · 被引用 378 次
相关 Paper
- Instance Localization for Self-Supervised Detection PretrainingCeyuan Yang, Zhirong Wu, Bolei Zhou, Stephen LinCVPR 2021
- Self-Supervised Transformers for Unsupervised Object Discovery using Normalized CutYangtao Wang, Xi Shen, Shell Xu Hu, Yuan Yuan 等CVPR 2022 · 被引用 143 次
- Finding Distributed Object-Centric Properties in Self-Supervised TransformersSamyak Rawlekar, Amitabh Swain, Yujun Cai, Yiwei Wang 等CVPR 2026 · 被引用 1 次
- Semantic-Aware Superpixel for Weakly Supervised Semantic SegmentationSangtae Kim, Daeyoung Park, Byonghyo ShimAAAI 2023 · 被引用 35 次
- Unsupervised Semantic Segmentation with Self-supervised Object-centric RepresentationsAndrii Zadaianchuk, Matthäus Kleindessner, Yi Zhu, Francesco Locatello 等ICLR 2023 · 被引用 16 次
