HASSOD: Hierarchical Adaptive Self-Supervised Object Detection
Shengcao Cao, Dhiraj Joshi, Liangyan Gui, Yu-Xiong Wang
Abstract
The human visual perception system demonstrates exceptional capabilities in learning without explicit supervision and understanding the part-to-whole composition of objects. Drawing inspiration from these two abilities, we propose Hierarchical Adaptive Self-Supervised Object Detection (HASSOD), a novel approach that learns to detect objects and understand their compositions without human supervision. HASSOD employs a hierarchical adaptive clustering strategy to group regions into object masks based on self-supervised visual representations, adaptively determining the number of objects per image. Furthermore, HASSOD identifies the hierarchical levels of objects in terms of composition, by analyzing coverage relations between masks and constructing tree structures. This additional self-supervised learning task leads to improved detection performance and enhanced interpretability. Lastly, we abandon the inefficient multi-round self-training process utilized in prior methods and instead adapt the Mean Teacher framework from semi-supervised learning, which leads to a smoother and more efficient training process. Through extensive experiments on prevalent image datasets, we demonstrate the superiority of HASSOD over existing methods, thereby advancing the state of the art in self-supervised object detection. Notably, we improve Mask AR from 20.2 to 22.5 on LVIS, and from 17.0 to 26.0 on SA-1B. Project page: https://HASSOD-NeurIPS23.github.io.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Segment Anything without SupervisionXudong Wang, Jingfeng Yang, Trevor DarrellNeurIPS 2024 · 36 citations
- DetKDS: Knowledge Distillation Search for Object DetectorsLujun Li, Yufan Bao, Peijie Dong, Chuanguang Yang et al.ICML 2024 · 35 citations
- DiPEx: Dispersing Prompt Expansion for Class-Agnostic Object DetectionJia Syuen Lim, Zhuoxiao Chen, Zhi Chen, Mahsa Baktashmotlagh et al.NeurIPS 2024 · 19 citations
- SOHES: Self-supervised Open-world Hierarchical Entity SegmentationShengcao Cao, Jiuxiang Gu, Jason Kuen, Hao Tan et al.ICLR 2024 · 3 citations
- Scene-Centric Unsupervised Panoptic SegmentationOliver Hahn, Christoph Reich, Nikita Araslanov, Daniel Cremers et al.CVPR 2025
Builds on10
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Objects365: A Large-Scale, High-Quality Dataset for Object DetectionShuai Shao, Zeming Li, Tianyuan Zhang, Chao Peng et al.ICCV 2019 · 1,018 citations
- Unbiased Teacher for Semi-Supervised Object DetectionYen-Cheng Liu, Chih-Yao Ma, Zijian He, Chia-Wen Kuo et al.ICLR 2021 · 603 citations
- Large-Scale Unsupervised Object DiscoveryHuy V. Vo, Elena Sizikova, Cordelia Schmid, Patrick Pérez et al.NeurIPS 2021 · 63 citations
Related papers
- Self-Supervised Object Detection from Egocentric VideosPeri Akiva, Jing Huang, Kevin J. Liang, Rama Kovvuri et al.ICCV 2023 · 9 citations
- Unsupervised Semantic Segmentation with Self-supervised Object-centric RepresentationsAndrii Zadaianchuk, Matthäus Kleindessner, Yi Zhu, Francesco Locatello et al.ICLR 2023 · 16 citations
- Adaptive Hierarchical Representation Learning for Long-Tailed Object DetectionBanghuai LiCVPR 2022 · 16 citations
- Self-Supervised Object Detection via Generative Image SynthesisSiva Karthik Mustikovela, Shalini De Mello, Aayush Prakash, Umar Iqbal et al.ICCV 2021 · 3 citations
- Open-Vocabulary Object Detection via Language HierarchyJiaxing Huang, Jingyi Zhang, Kai Jiang, Shijian LuNeurIPS 2024 · 16 citations
