Capturing Head Avatar with Hand Contacts from a Monocular Video
Haonan He, Yufeng Zheng, Jie Song
摘要
Photorealistic 3D head avatars are vital for telepresence, gaming, and VR. However, most methods focus solely on facial regions, ignoring natural hand-face interactions, such as a hand resting on the chin or fingers gently touching the cheek, which convey cognitive states like pondering. In this work, we present a novel framework that jointly learns detailed head avatars and the non-rigid deformations induced by hand-face interactions. There are two principal challenges in this task. First, naively tracking hand and face separately fails to capture their relative poses. To overcome this, we propose to combine depth order loss with contact regularization during pose tracking, ensuring correct spatial relationships between the face and hand. Second, no publicly available priors exist for hand-induced deformations, making them non-trivial to learn from monocular videos. To address this, we learn a PCA basis specific to hand-induced facial deformations from a face-hand interaction dataset. This reduces the problem to estimating a compact set of PCA parameters rather than a full spatial deformation field. Furthermore, inspired by physics-based simulation, we incorporate a contact loss that provides additional supervision, significantly reducing interpenetration artifacts and enhancing the physical plausibility of the results. We evaluate our approach on RGB(D) videos captured by an iPhone. Additionally, to better evaluate the reconstructed geometry, we construct a synthetic dataset of avatars with various types of hand interactions. We show that our method can capture better appearance and more accurate deforming geometry of the face than SOTA surface reconstruction methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Multiview Neural Surface Reconstruction by Disentangling Geometry and AppearanceLior Yariv, Yoni Kasten, Dror Moran, Meirav Galun 等NeurIPS 2020 · 被引用 1,010 次
- Learning an animatable detailed 3D face model from in-the-wild imagesYao Feng, Haiwen Feng, Michael J. Black, Timo BolkartSIGGRAPH 2021 · 被引用 662 次
- GaussianAvatars: Photorealistic Head Avatars with Rigged 3D GaussiansShenhan Qian, Tobias Kirschstein, Liam Schoneveld, Davide Davoli 等CVPR 2024 · 被引用 175 次
- Reconstructing Hands in 3D with TransformersGeorgios Pavlakos, Dandan Shan, Ilija Radosavovic, Angjoo Kanazawa 等CVPR 2024 · 被引用 110 次
- MonoGaussianAvatar: Monocular Gaussian Point-based Head AvatarYufan Chen, Lizhen Wang, Qijing Li, Hongjiang Xiao 等SIGGRAPH 2024 · 被引用 85 次
相关 Paper
- InteractAvatar: Modeling Hand-Face Interaction in Photorealistic Avatars with Deformable GaussiansKefan Chen, Sreyas Mohan, Justin Theiss, Sergiu Oprea 等ICCV 2025 · 被引用 2 次
- The eyes have it: an integrated eye and face model for photorealistic facial animationGabriel Schwartz, Shih-En Wei, Te-Li Wang, Stephen Lombardi 等SIGGRAPH 2020 · 被引用 54 次
- InterAnimate: Taming Region-Aware Diffusion Model for Realistic Human Interaction AnimationYukang Lin, Yan Hong, Zunnan Xu, Xindi Li 等ACM MM 2025
- Physics-Aware Hand-Object Interaction DenoisingHaowen Luo, Yunze Liu, Li YiCVPR 2024 · 被引用 2 次
- PhysHanDI: Physics-Based Reconstruction of Hand-Deformable Object InteractionsJihyun Lee, Changmin Lee, Donghwan Kim, Tae-Kyun KimICML 2026
