PoseKernelLifter: Metric Lifting of 3D Human Pose using Sound
Zhijian Yang, Xiaoran Fan, Volkan Isler, Hyunsoo Park
摘要
Reconstructing the 3D pose of a person in metric scale from a single view image is a geometrically ill-posed problem. For example, we can not measure the exact distance of a person to the camera from a single view image without additional scene assumptions (e.g., known height). Existing learning based approaches circumvent this issue by reconstructing the 3D pose up to scale. However, there are many applications such as virtual telepresence, robotics, and augmented reality that require metric scale reconstruction. In this paper, we show that audio signals recorded along with an image, provide complementary information to reconstruct the metric 3D pose of the person. The key insight is that as the audio signals traverse across the 3D space, their interactions with the body provide metric information about the body's pose. Based on this insight, we introduce a time-invariant transfer function called pose kernel -- the impulse response of audio signals induced by the body pose. The main properties of the pose kernel are that (1) its envelope highly correlates with 3D pose, (2) the time response corresponds to arrival time, indicating the metric distance to the microphone, and (3) it is invariant to changes in the scene geometry configurations. Therefore, it is readily generalizable to unseen scenes. We design a multi-stage 3D CNN that fuses audio and visual signals and learns to reconstruct 3D pose in a metric scale. We show that our multi-modal method produces accurate metric reconstruction in real world scenes, which is not possible with state-of-the-art lifting approaches including parametric mesh regression and depth regression.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper18
- Exploiting Spatial-Temporal Relationships for 3D Pose Estimation via Graph Convolutional NetworksYujun Cai, Liuhao Ge, Jun Liu, Jianfei Cai 等ICCV 2019 · 被引用 504 次
- Learnable Triangulation of Human PoseKarim Iskakov, Egor Burkov, Victor S. Lempitsky, Yury MalkovICCV 2019 · 被引用 419 次
- Towards 3D human pose construction using wifiWenjun Jiang, Hongfei Xue, Chenglin Miao, Shiyang Wang 等MobiCom 2020 · 被引用 282 次
- Optimizing Network Structure for 3D Human Pose EstimationHai Ci, Chunyu Wang, Xiaoxuan Ma, Yizhou WangICCV 2019 · 被引用 267 次
- XNect: real-time multi-person 3D motion capture with a single RGB cameraDushyant Mehta, Oleksandr Sotnychenko, Franziska Mueller, Weipeng Xu 等SIGGRAPH 2020 · 被引用 267 次
相关 Paper
- Learning Human Mesh Recovery in 3D ScenesZehong Shen, Zhi Cen, Sida Peng, Qing Shuai 等CVPR 2023
- Camera Distance-Aware Top-Down Approach for 3D Multi-Person Pose Estimation From a Single RGB ImageGyeongsik Moon, Ju Yong Chang, Kyoung Mu LeeICCV 2019 · 被引用 368 次
- Towards Alleviating the Modeling Ambiguity of Unsupervised Monocular 3D Human Pose EstimationZhenbo Yu, Bingbing Ni, Jingwei Xu, Junjie Wang 等ICCV 2021 · 被引用 39 次
- MetricHMSR: Metric Human Mesh and Scene Recovery from Monocular ImagesChentao Song, He Zhang, Haolei Yuan, Haozhe Lin 等CVPR 2026 · 被引用 5 次
- Implicit 3D Human Mesh Recovery using Consistency with Pose and Shape from Unseen-viewHanbyel Cho, Yooshin Cho, Jaesung Ahn, Junmo KimCVPR 2023
