xMHashSeg: Cross-modal Hash Learning for Training-free Unsupervised LiDAR Semantic Segmentation
Jialong Zhang, Yachao Zhang, Yao Wu, Jiangming Shi, Fangyong Wang, Yanyun Qu
Abstract
3D semantic segmentation serves as a fundamental component in many applications, such as autonomous driving and medical image analysis. Although recent methods have advanced the field, adapting these methods to new environments or object categories without extensive retraining remains a significant challenge. To address this, we introduce xMHashSeg, a novel training-free cross-modal Li-DAR semantic segmentation framework. xMHashSeg leverages foundation models and non-parametric network to extract features from 2D images and 3D point clouds, subsequently integrating these features through hash learning. Specifically, We develop point-SANN, a novel self-adaption non-parametric network that can extract robust 3D features from raw point clouds, while 2D features are directly extracted through the foundation model DINOv2. To reconcile inconsistencies across different modals, we introduce a Hash Code Learning Module that projects all information into a common hash space, learning a consistent hash code that enhances feature integration. Additionally, depth maps are utilized as an intermediary form between 2D and 3D data to facilitate convergence during hash code learning. Our experimental results on various multi-modality datasets demonstrate that xMHashSeg outperforms zero-shot learning approaches and achieve performance close to that of unsupervised domain adaptation and test-time adaptation methods, without requiring any annotations or additional training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fdd82bdf-9248-4ff9-ac2a-4ae7096edefbBuilds on16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR SequencesJens Behley, Martin Garbade, Andres Milioto, Jan Quenzel et al.ICCV 2019 · 2,345 citations
- Perturbed Self-Distillation: Weakly Supervised Large-Scale Point Cloud Semantic SegmentationYachao Zhang, Yanyun Qu, Yuan Xie, Zonghao Li et al.ICCV 2021 · 138 citations
- Sparse-to-dense Feature Matching: Intra and Inter domain Cross-modal Learning in Domain Adaptation for 3D Semantic SegmentationDuo Peng, Yinjie Lei, Wen Li, Pingping Zhang et al.ICCV 2021 · 79 citations
Related papers
- Cross-modal & Cross-domain Learning for Unsupervised LiDAR Semantic SegmentationYiyang Chen, Shanshan Zhao, Changxing Ding, Liyao Tang et al.ACM MM 2023 · 4 citations
- Cross-Modal Contrastive Learning for Domain Adaptation in 3D Semantic SegmentationBowei Xing, Xianghua Ying, Ruibin Wang, Jinfa Yang et al.AAAI 2023 · 23 citations
- xMUDA: Cross-Modal Unsupervised Domain Adaptation for 3D Semantic SegmentationMaximilian Jaritz, Tuan-Hung Vu, Raoul de Charette, Émilie Wirbel et al.CVPR 2020
- CrossPoint: Self-Supervised Cross-Modal Contrastive Learning for 3D Point Cloud UnderstandingMohamed Afham, Isuru Dissanayake, Dinithi Dissanayake, Amaya Dharmasiri et al.CVPR 2022 · 286 citations
- PointDC: Unsupervised Semantic Segmentation of 3D Point Clouds via Cross-modal Distillation and Super-Voxel ClusteringZisheng Chen, Hongbin Xu, Weitao Chen, Zhipeng Zhou et al.ICCV 2023 · 21 citations
