HUNTER: Unsupervised Human-Centric 3D Detection via Transferring Knowledge from Synthetic Instances to Real Scenes
Yichen Yao, Zimo Jiang, Yujing Sun, Zhencai Zhu, Xinge Zhu, Runnan Chen, Yuexin Ma
Abstract
Human-centric 3D scene understanding has recently drawn increasing attention, driven by its critical impact on robotics. However, human-centric real-life scenarios are extremely diverse and complicated, and humans have intri-cate motions and interactions. With limited labeled data, supervised methods are difficult to generalize to general scenarios, hindering real-life applications. Mimicking human intelligence, we propose an unsupervised 3D detection method for human-centric scenarios by transferring the knowledge from synthetic human instances to real scenes. To bridge the gap between the distinct data representations and feature distributions of synthetic models and real point clouds, we introduce novel modules for effective instance-to-scene representation transfer and synthetic-to-real feature alignment. Remarkably, our method exhibits superior performance compared to current state-of-the-art techniques, achieving 87.8% improvement in mAP and closely approaching the performance of fully supervised methods (62.15 mAP vs. 69.02 mAP) on HuCenLife Dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Gait Recognition in Large-scale Free Environment via Single LiDARXiao Han, Yiming Ren, Peishan Cong, Yujing Sun et al.ACM MM 2024 · 10 citations
- Towards Practical Human Motion Prediction with LiDAR Point CloudsXiao Han, Yiming Ren, Yichen Yao, Yujing Sun et al.ACM MM 2024 · 2 citations
- AnnofreeOD: Detecting All Classes at Low Frame Rates Without Human AnnotationsBoyi Sun, Yuhang Liu, Houxin He, Yonglin Tian et al.ICCV 2025 · 1 citation
Builds on23
- SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR SequencesJens Behley, Martin Garbade, Andres Milioto, Jan Quenzel et al.ICCV 2019 · 2,345 citations
- Segment Any Point Cloud Sequences by Distilling Vision Foundation ModelsYouquan Liu, Lingdong Kong, Jun Cen, Runnan Chen et al.NeurIPS 2023 · 169 citations
- Self-Supervised Transformers for Unsupervised Object Discovery using Normalized CutYangtao Wang, Xi Shen, Shell Xu Hu, Yuan Yuan et al.CVPR 2022 · 143 citations
- Towards Label-free Scene Understanding by Vision Foundation ModelsRunnan Chen, Youquan Liu, Lingdong Kong, Nenglun Chen et al.NeurIPS 2023 · 82 citations
- RandomRooms: Unsupervised Pre-training from Synthetic Shapes and Randomized Layouts for 3D Object DetectionYongming Rao, Benlin Liu, Yi Wei, Jiwen Lu et al.ICCV 2021 · 58 citations
Related papers
- Human-centric Scene Understanding for 3D Large-scale ScenariosYiteng Xu, Peishan Cong, Yichen Yao, Runnan Chen et al.ICCV 2023 · 34 citations
- UniDet3D: Multi-dataset Indoor 3D Object DetectionMaksim Kolodiazhnyi, Anna Vorontsova, Matvey Skripkin, Danila Rukhovich et al.AAAI 2025 · 7 citations
- 3D Segmentation of Humans in Point Clouds with Synthetic DataAyça Takmaz, Jonas Schult, Irem Kaftan, Mertcan Akçay et al.ICCV 2023 · 31 citations
- Self-Supervised Pretraining for Large-Scale Point CloudsZaiwei Zhang, Min Bai, Li Erran LiNeurIPS 2022 · 12 citations
- A Unified Framework for Human-centric Point Cloud Video UnderstandingYiteng Xu, Kecheng Ye, Xiao Han, Yiming Ren et al.CVPR 2024 · 4 citations
