Human-centric Scene Understanding for 3D Large-scale Scenarios
Yiteng Xu, Peishan Cong, Yichen Yao, Runnan Chen, Yuenan Hou, Xinge Zhu, Xuming He, Jingyi Yu, Yuexin Ma
Abstract
Human-centric scene understanding is significant for real-world applications, but it is extremely challenging due to the existence of diverse human poses and actions, complex human-environment interactions, severe occlusions in crowds, etc. In this paper, we present a large-scale multi-modal dataset for human-centric scene under-standing, dubbed HuCenLife, which is collected in diverse daily-life scenarios with rich and fine-grained annotations. Our HuCenLife can benefit many 3D perception tasks, such as segmentation, detection, action recognition, etc., and we also provide benchmarks for these tasks to facilitate related research. In addition, we design novel modules for LiDAR-based segmentation and action recognition, which are more applicable for large-scale human-centric scenarios and achieve state-of-the-art performance. The dataset and code can be found at https://github.com/4DVLab/HuCenLife.git.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- UniSeg: A Unified Multi-Modal LiDAR Segmentation Network and the OpenPCSeg CodebaseYouquan Liu, Runnan Chen, Xin Li, Lingdong Kong et al.ICCV 2023 · 94 citations
- Towards Label-free Scene Understanding by Vision Foundation ModelsRunnan Chen, Youquan Liu, Lingdong Kong, Nenglun Chen et al.NeurIPS 2023 · 82 citations
- LiveHPS: LiDAR-Based Scene-Level Human Pose and Shape Estimation in Free EnvironmentYiming Ren, Xiao Han, Chengfeng Zhao, Jingya Wang et al.CVPR 2024 · 14 citations
- Gait Recognition in Large-scale Free Environment via Single LiDARXiao Han, Yiming Ren, Peishan Cong, Yujing Sun et al.ACM MM 2024 · 10 citations
- Bridging Language and Geometric Primitives for Zero-shot Point Cloud SegmentationRunnan Chen, Xinge Zhu, Nenglun Chen, Wei Li et al.ACM MM 2023 · 8 citations
Builds on41
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- KPConv: Flexible and Deformable Convolution for Point CloudsHugues Thomas, Charles R. Qi, Jean-Emmanuel Deschaud, Beatriz Marcotegui et al.ICCV 2019 · 3,193 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR SequencesJens Behley, Martin Garbade, Andres Milioto, Jan Quenzel et al.ICCV 2019 · 2,345 citations
Related papers
- HUNTER: Unsupervised Human-Centric 3D Detection via Transferring Knowledge from Synthetic Instances to Real ScenesYichen Yao, Zimo Jiang, Yujing Sun, Zhencai Zhu et al.CVPR 2024 · 4 citations
- UAVScenes: A Multi-Modal Dataset for UAVsSijie Wang, Siqi Li, Yawei Zhang, Shangshu Yu et al.ICCV 2025 · 9 citations
- STCrowd: A Multimodal Dataset for Pedestrian Perception in Crowded ScenesPeishan Cong, Xinge Zhu, Feng Qiao, Yiming Ren et al.CVPR 2022 · 43 citations
- HmPEAR: A Dataset for Human Pose Estimation and Action RecognitionYitai Lin, Zhijie Wei, Wanfa Zhang, Xiping Lin et al.ACM MM 2024 · 5 citations
- EmbodiedScan: A Holistic Multi-Modal 3D Perception Suite Towards Embodied AITai Wang, Xiaohan Mao, Chenming Zhu, Runsen Xu et al.CVPR 2024 · 53 citations
