Unsupervised 3D Perception with 2D Vision-Language Distillation for Autonomous Driving
Mahyar Najibi, Jingwei Ji, Yin Zhou, Charles R. Qi, Xinchen Yan, Scott Ettinger, Dragomir Anguelov
Abstract
Closed-set 3D perception models trained on only a predefined set of object categories can be inadequate for safety critical applications such as autonomous driving where new object types can be encountered after deployment. In this paper, we present a multi-modal auto labeling pipeline capable of generating amodal 3D bounding boxes and tracklets for training models on open-set categories without 3D human labels. Our pipeline exploits motion cues inherent in point cloud sequences in combination with the freely available 2D image-text pairs to identify and track all traffic participants. Compared to the recent studies in this domain, which can only provide class-agnostic auto labels limited to moving objects, our method can handle both static and moving objects in the unsupervised manner and is able to output open-vocabulary semantic labels thanks to the proposed vision-language knowledge distillation. Experiments on the Waymo Open Dataset show that our approach outperforms the prior work by significant margins on various unsupervised 3D perception tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d7578b5d-2eb6-4a92-9d73-ab8e79127909Cited by top-tier papers13
- RefAV: Towards Planning-Centric Scenario MiningCainan Davidson, Deva Ramanan, Neehar PeriCVPR 2026 · 17 citations
- OpenBox: Annotate Any Bounding Boxes in 3DIn-Jae Lee, Mungyeom Kim, Kwonyoung Ryu, Pierre Musacchio et al.NeurIPS 2025 · 7 citations
- Hawaii: Hierarchical Visual Knowledge Transfer for Efficient Vision-Language ModelsYimu Wang, Mozhgan Nasr Azadani, Sean Sedwards, Krzysztof CzarneckiNeurIPS 2025 · 6 citations
- HUNTER: Unsupervised Human-Centric 3D Detection via Transferring Knowledge from Synthetic Instances to Real ScenesYichen Yao, Zimo Jiang, Yujing Sun, Zhencai Zhu et al.CVPR 2024 · 4 citations
- OV-SCAN: Semantically Consistent Alignment for Novel Object Discovery in Open-Vocabulary 3D Object DetectionAdrian Chow, Evelien Riddell, Yimu Wang, Sean Sedwards et al.ICCV 2025 · 2 citations
Builds on29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Deep Hough Voting for 3D Object Detection in Point CloudsCharles R. Qi, Or Litany, Kaiming He, Leonidas J. GuibasICCV 2019 · 1,467 citations
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 1,274 citations
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve et al.ICCV 2021 · 1,114 citations
Related papers
- Offboard 3D Object Detection From Point Cloud SequencesCharles R. Qi, Yin Zhou, Mahyar Najibi, Pei Sun et al.CVPR 2021
- ZOPP: A Framework of Zero-shot Offboard Panoptic Perception for Autonomous DrivingTao Ma, Hongbin Zhou, Qiusheng Huang, Xuemeng Yang et al.NeurIPS 2024 · 8 citations
- OVTrack: Open-Vocabulary Multiple Object TrackingSiyuan Li, Tobias Fischer, Lei Ke, Henghui Ding et al.CVPR 2023
- 3D Annotation-Free Learning by Distilling 2D Open-Vocabulary Segmentation Models for Autonomous DrivingBoyi Sun, Yuhang Liu, Xingxia Wang, Bin Tian et al.AAAI 2025 · 6 citations
- Unsupervised Object Detection With LIDAR CluesHao Tian, Yuntao Chen, Jifeng Dai, Zhaoxiang Zhang et al.CVPR 2021
