Human-centric Scene Understanding for 3D Large-scale Scenarios
Yiteng Xu, Peishan Cong, Yichen Yao, Runnan Chen, Yuenan Hou, Xinge Zhu, Xuming He, Jingyi Yu, Yuexin Ma
摘要
Human-centric scene understanding is significant for real-world applications, but it is extremely challenging due to the existence of diverse human poses and actions, complex human-environment interactions, severe occlusions in crowds, etc. In this paper, we present a large-scale multi-modal dataset for human-centric scene under-standing, dubbed HuCenLife, which is collected in diverse daily-life scenarios with rich and fine-grained annotations. Our HuCenLife can benefit many 3D perception tasks, such as segmentation, detection, action recognition, etc., and we also provide benchmarks for these tasks to facilitate related research. In addition, we design novel modules for LiDAR-based segmentation and action recognition, which are more applicable for large-scale human-centric scenarios and achieve state-of-the-art performance. The dataset and code can be found at https://github.com/4DVLab/HuCenLife.git.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- UniSeg: A Unified Multi-Modal LiDAR Segmentation Network and the OpenPCSeg CodebaseYouquan Liu, Runnan Chen, Xin Li, Lingdong Kong 等ICCV 2023 · 被引用 94 次
- Towards Label-free Scene Understanding by Vision Foundation ModelsRunnan Chen, Youquan Liu, Lingdong Kong, Nenglun Chen 等NeurIPS 2023 · 被引用 82 次
- LiveHPS: LiDAR-Based Scene-Level Human Pose and Shape Estimation in Free EnvironmentYiming Ren, Xiao Han, Chengfeng Zhao, Jingya Wang 等CVPR 2024 · 被引用 14 次
- Gait Recognition in Large-scale Free Environment via Single LiDARXiao Han, Yiming Ren, Peishan Cong, Yujing Sun 等ACM MM 2024 · 被引用 10 次
- Bridging Language and Geometric Primitives for Zero-shot Point Cloud SegmentationRunnan Chen, Xinge Zhu, Nenglun Chen, Wei Li 等ACM MM 2023 · 被引用 8 次
它引用的顶会 Paper41
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- KPConv: Flexible and Deformable Convolution for Point CloudsHugues Thomas, Charles R. Qi, Jean-Emmanuel Deschaud, Beatriz Marcotegui 等ICCV 2019 · 被引用 3,193 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
- SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR SequencesJens Behley, Martin Garbade, Andres Milioto, Jan Quenzel 等ICCV 2019 · 被引用 2,345 次
相关 Paper
- HUNTER: Unsupervised Human-Centric 3D Detection via Transferring Knowledge from Synthetic Instances to Real ScenesYichen Yao, Zimo Jiang, Yujing Sun, Zhencai Zhu 等CVPR 2024 · 被引用 4 次
- UAVScenes: A Multi-Modal Dataset for UAVsSijie Wang, Siqi Li, Yawei Zhang, Shangshu Yu 等ICCV 2025 · 被引用 9 次
- STCrowd: A Multimodal Dataset for Pedestrian Perception in Crowded ScenesPeishan Cong, Xinge Zhu, Feng Qiao, Yiming Ren 等CVPR 2022 · 被引用 43 次
- HmPEAR: A Dataset for Human Pose Estimation and Action RecognitionYitai Lin, Zhijie Wei, Wanfa Zhang, Xiping Lin 等ACM MM 2024 · 被引用 5 次
- EmbodiedScan: A Holistic Multi-Modal 3D Perception Suite Towards Embodied AITai Wang, Xiaohan Mao, Chenming Zhu, Runsen Xu 等CVPR 2024 · 被引用 53 次
