Motal: Unsupervised 3D Object Detection by Modality and Task-Specific Knowledge Transfer
Hai Wu, Hongwei Lin, Xusheng Guo, Xin Li, Mingming Wang, Cheng Wang, Chenglu Wen
Abstract
The performance of unsupervised 3D object classification and bounding box regression relies heavily on the quality of initial pseudo-labels. Traditionally, the labels of classification and regression are represented by a single set of candidate boxes generated by motion or geometry heuristics. However, due to the similarity of many objects to the background in shape or lack of motion, the labels often fail to achieve high accuracy in two tasks simultaneously. Using these labels to directly train the network results in decreased detection performance. To address this challenge, we introduce Motal that performs unsupervised 3D object detection by Modality and task-specific knowledge transfer. Motal decouples the pseudo-labels into two sets of candidates, from which Motal discovers classification knowledge by motion and image appearance prior, and discovers box regression knowledge by geometry prior, respectively. Motal finally transfers all knowledge to a single student network by a TMT (Task-specific Masked Training) scheme, attaining high performance in both classification and regression. Motal can greatly enhance various unsupervised methods by about 2× mAP. For example, on the WOD test set, Motal improves the state-of-the-art CPD by 21.56% mAP L1 (from 20.54% to 42.10%) and 19.90% mAP L2 (from 18.18% to 38.08%). These achievements highlight the significance of our method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 842ec7a3-593a-4dd8-bd99-160579d2ac3bCited by top-tier papers1
Ask how each one uses itBuilds on21
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Voxel R-CNN: Towards High Performance Voxel-based 3D Object DetectionJiajun Deng, Shaoshuai Shi, Peiwei Li, Wengang Zhou et al.AAAI 2021 · 1,128 citations
- Sparse Fuse Dense: Towards High Quality 3D Detection with Depth CompletionXiaopei Wu, Liang Peng, Honghui Yang, Liang Xie et al.CVPR 2022 · 248 citations
- UniTR: A Unified and Efficient Multi-Modal Transformer for Bird's-Eye-View RepresentationHaiyang Wang, Hao Tang, Shaoshuai Shi, Aoxue Li et al.ICCV 2023 · 106 citations
Related papers
- Rethinking the Route Towards Weakly Supervised Object LocalizationChen-Lin Zhang, Yun-Hao Cao, Jianxin WuCVPR 2020
- Transferable Semi-Supervised 3D Object Detection From RGB-D DataYew Siang Tang, Gim Hee LeeICCV 2019 · 41 citations
- Unifying Voxel-based Representation with Transformer for 3D Object DetectionYanwei Li, Yilun Chen, Xiaojuan Qi, Zeming Li et al.NeurIPS 2022 · 401 citations
- mDALU: Multi-Source Domain Adaptation and Label Unification with Partial DatasetsRui Gong, Dengxin Dai, Yuhua Chen, Wen Li et al.ICCV 2021 · 27 citations
- Anomaly Detection in Video via Self-Supervised and Multi-Task LearningMariana-Iuliana Georgescu, Antonio Barbalau, Radu Tudor Ionescu, Fahad Shahbaz Khan et al.CVPR 2021
