Pre-training a Density-Aware Pose Transformer for Robust LiDAR-based 3D Human Pose Estimation
Xiaoqi An, Lin Zhao, Chen Gong, Jun Li, Jian Yang
摘要
With the rapid development of autonomous driving, LiDAR-based 3D Human Pose Estimation (3D HPE) is becoming a research focus. However, due to the noise and sparsity of LiDAR-captured point clouds, robust human pose estimation remains challenging. Most of the existing methods use temporal information, multi-modal fusion, or SMPL optimization to correct biased results. In this work, we try to obtain sufficient information for 3D HPE only by modeling the intrinsic properties of low-quality point clouds. Hence, a simple yet powerful method is proposed, which provides insights both on modeling and augmentation of point clouds. Specifically, we first propose a concise and effective density-aware pose transformer (DAPT) to get stable keypoint representations. By using a set of joint anchors and a carefully designed exchange module, valid information is extracted from point clouds with different densities. Then 1D heatmaps are utilized to represent the precise locations of the keypoints. Secondly, a comprehensive LiDAR human synthesis and augmentation method is proposed to pre-train the model, enabling it to acquire a better human body prior. We increase the diversity of point clouds by randomly sampling human positions and orientations and by simulating occlusions through the addition of laser-level masks. Extensive experiments have been conducted on multiple datasets, including IMU-annotated LidarHuman26M, SLOPER4D, and manually annotated Waymo Open Dataset v2.0 (Waymo), HumanM3. Our method demonstrates SOTA performance in all scenarios. In particular, compared with LPFormer on Waymo, we reduce the average MPJPE by 10.0mm. Compared with PRN on SLOPER4D, we notably reduce the average MPJPE by 20.7mm.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- VIMCAN: Visual-Inertial 3D Human Pose Estimation with Hybrid Mamba-Cross-Attention NetworkZepeng Yang, Junxuan Bai, Hao Li, Ju Dai 等CVPR 2026 · 被引用 1 次
- Bézier Degradation Modeling for LiDAR-based Human Motion CaptureXiaoqi An, Lin Zhao, Jun Li, Chen Gong 等CVPR 2026
- Shaping Without Tearing: Controllable Diffeomorphic Deformations for Topology-Preserving 3D Point Cloud AugmentationJian Bi, Qianliang Wu, Jianjun Qian, Lei Luo 等AAAI 2026
它引用的顶会 Paper20
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll 等ICCV 2019 · 被引用 1,784 次
- ViTPose: Simple Vision Transformer Baselines for Human Pose EstimationYufei Xu, Jing Zhang, Qiming Zhang, Dacheng TaoNeurIPS 2022 · 被引用 1,105 次
- Learnable Triangulation of Human PoseKarim Iskakov, Egor Burkov, Victor S. Lempitsky, Yury MalkovICCV 2019 · 被引用 419 次
- Unsupervised Point Cloud Pre-training via Occlusion CompletionHanchen Wang, Qi Liu, Xiangyu Yue, Joan Lasenby 等ICCV 2021 · 被引用 323 次
- Online Knowledge Distillation for Efficient Pose EstimationZheng Li, Jingwen Ye, Mingli Song, Ying Huang 等ICCV 2021 · 被引用 123 次
相关 Paper
- Neighborhood-Enhanced 3D Human Pose Estimation with Monocular LiDAR in Long-Range Outdoor ScenesJingyi Zhang, Qihong Mao, Guosheng Hu, Siqi Shen 等AAAI 2024 · 被引用 11 次
- Towards Practical Human Motion Prediction with LiDAR Point CloudsXiao Han, Yiming Ren, Yichen Yao, Yujing Sun 等ACM MM 2024 · 被引用 2 次
- LiDARCap: Long-range Markerless 3D Human Motion Capture with LiDAR Point CloudsJialian Li, Jingyi Zhang, Zhiyong Wang, Siqi Shen 等CVPR 2022 · 被引用 49 次
- MSeg3D: Multi-Modal 3D Semantic Segmentation for Autonomous DrivingJiale Li, Hang Dai, Hao Han, Yong DingCVPR 2023
- Embracing Single Stride 3D Object Detector with Sparse TransformerLue Fan, Ziqi Pang, Tianyuan Zhang, Yu-Xiong Wang 等CVPR 2022
