4D Unsupervised Object Discovery
Yuqi Wang, Yuntao Chen, Zhaoxiang Zhang
Abstract
Object discovery is a core task in computer vision. While fast progresses have been made in supervised object detection, its unsupervised counterpart remains largely unexplored. With the growth of data volume, the expensive cost of annotations is the major limitation hindering further study. Therefore, discovering objects without annotations has great significance. However, this task seems impractical on still-image or point cloud alone due to the lack of discriminative information. Previous studies underlook the crucial temporal information and constraints naturally behind multi-modal inputs. In this paper, we propose 4D unsupervised object discovery, jointly discovering objects from 4D data -- 3D point clouds and 2D RGB images with temporal information. We present the first practical approach for this task by proposing a ClusterNet on 3D point clouds, which is jointly iteratively optimized with a 2D localization network. Extensive experiments on the large-scale Waymo Open Dataset suggest that the localization network and ClusterNet achieve competitive performance on both class-agnostic 2D object detection and 3D instance segmentation, bridging the gap between unsupervised methods and full supervised ones. Codes and models will be made available at https://github.com/Robertwyq/LSMOL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- UNION: Unsupervised 3D Object Detection using Object Appearance-based Pseudo-ClassesTed de Vries Lentsch, Holger Caesar, Dariu GavrilaNeurIPS 2024 · 30 citations
- HUNTER: Unsupervised Human-Centric 3D Detection via Transferring Knowledge from Synthetic Instances to Real ScenesYichen Yao, Zimo Jiang, Yujing Sun, Zhencai Zhu et al.CVPR 2024 · 4 citations
- MonoSOWA: Scalable Monocular 3D Object Detector Without Human AnnotationsJan Skvrna, Lukás NeumannICCV 2025 · 3 citations
Builds on14
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
- Deep Hough Voting for 3D Object Detection in Point CloudsCharles R. Qi, Or Litany, Kaiming He, Leonidas J. GuibasICCV 2019 · 1,467 citations
- Neural Scene Flow PriorXueqian Li, Jhony Kaesemodel Pontes, Simon LuceyNeurIPS 2021 · 136 citations
Related papers
- 4D-Net for Learned Multi-Modal AlignmentA. J. Piergiovanni, Vincent Casser, Michael S. Ryoo, Anelia AngelovaICCV 2021 · 69 citations
- Unsupervised Object Detection With LIDAR CluesHao Tian, Yuntao Chen, Jifeng Dai, Zhaoxiang Zhang et al.CVPR 2021
- OGC: Unsupervised 3D Object Segmentation from Rigid Dynamics of Point CloudsZiyang Song, Bo YangNeurIPS 2022 · 41 citations
- Weakly Supervised 3D Object Detection from Point CloudsZengyi Qin, Jinglu Wang, Yan LuACM MM 2020 · 68 citations
- PointDC: Unsupervised Semantic Segmentation of 3D Point Clouds via Cross-modal Distillation and Super-Voxel ClusteringZisheng Chen, Hongbin Xu, Weitao Chen, Zhipeng Zhou et al.ICCV 2023 · 21 citations
