3D Annotation-Free Learning by Distilling 2D Open-Vocabulary Segmentation Models for Autonomous Driving
Boyi Sun, Yuhang Liu, Xingxia Wang, Bin Tian, Long Chen, Fei-Yue Wang
Abstract
Point cloud data labeling is considered a time-consuming and expensive task in autonomous driving, whereas annotationfree learning training can avoid it by learning point cloud representations from unannotated data. In this paper, we propose AFOV, a novel 3D Annotation-Free framework assisted by 2D Open-Vocabulary segmentation models. It consists of two stages: In the first stage, we innovatively integrate highquality textual and image features of 2D open-vocabulary models and propose the Tri-Modal contrastive Pre-training (TMP). In the second stage, spatial mapping between point clouds and images is utilized to generate pseudo-labels, enabling cross-modal knowledge distillation. Besides, we introduce the Approximate Flat Interaction (AFI) to address the noise during alignment and label confusion. To validate the superiority of AFOV, extensive experiments are conducted on multiple related datasets. We achieved a recordbreaking 47.73% mIoU on the annotation-free 3D segmentation task in nuScenes, surpassing the previous best model by 3.13% mIoU. Meanwhile, the performance of fine-tuning with 1% data on nuScenes and SemanticKITTI reached a remarkable 51.75% mIoU and 48.14% mIoU, outperforming all previous pre-trained models. Our code is available at: https://github.com/sbysbysbys/AFOV
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4e027719-8c5f-417f-adfc-1ea141cd442cCited by top-tier papers2
- AnnofreeOD: Detecting All Classes at Low Frame Rates Without Human AnnotationsBoyi Sun, Yuhang Liu, Houxin He, Yonglin Tian et al.ICCV 2025 · 1 citation
- GeoPurify: A Data-Efficient Geometric Distillation Framework for Open-Vocabulary 3D SegmentationWeijia Dou, Xu Zhang, Yi Bin, Jian Liu et al.ICLR 2026 · 1 citation
Builds on22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR SequencesJens Behley, Martin Garbade, Andres Milioto, Jan Quenzel et al.ICCV 2019 · 2,345 citations
- Self-Supervised Pretraining of 3D Features on any Point-CloudZaiwei Zhang, Rohit Girdhar, Armand Joulin, Ishan MisraICCV 2021 · 333 citations
- Decoupling Zero-Shot Semantic SegmentationJian Ding, Nan Xue, Gui-Song Xia, Dengxin DaiCVPR 2022 · 255 citations
Related papers
- CLIP2Scene: Towards Label-efficient 3D Scene Understanding by CLIPRunnan Chen, Youquan Liu, Lingdong Kong, Xinge Zhu et al.CVPR 2023
- Transferring CLIP's Knowledge into Zero-Shot Point Cloud Semantic SegmentationYuanbin Wang, Shaofei Huang, Yulu Gao, Zhen Wang et al.ACM MM 2023 · 17 citations
- Open-Vocabulary Point-Cloud Object Detection without 3D AnnotationYuheng Lu, Chenfeng Xu, Xiaobao Wei, Xiaodong Xie et al.CVPR 2023
- MWSIS: Multimodal Weakly Supervised Instance Segmentation with 2D Box Annotations for Autonomous DrivingGuangfeng Jiang, Jun Liu, Yuzhi Wu, Wenlong Liao et al.AAAI 2024 · 11 citations
- OpenBox: Annotate Any Bounding Boxes in 3DIn-Jae Lee, Mungyeom Kim, Kwonyoung Ryu, Pierre Musacchio et al.NeurIPS 2025 · 7 citations
