Robust Egocentric Photo-realistic Facial Expression Transfer for Virtual Reality
Amin Jourabloo, Fernando De la Torre, Jason M. Saragih, Shih-En Wei, Stephen Lombardi, Te-Li Wang, Danielle Belko, Autumn Trimble, Hernán Badino
Abstract
Social presence, the feeling of being there with a “real” person, will fuel the next generation of communication systems driven by digital humans in virtual reality (VR). The best 3D video-realistic VR avatars that minimize the uncanny effect rely on person-specific (PS) models. However, these PS models are time-consuming to build and are typically trained with limited data variability, which results in poor generalization and robustness. Major sources of variability that affects the accuracy of facial expression transfer algorithms include using different VR headsets (e.g., camera configuration, slop of the headset), facial appearance changes over time (e.g., beard, make-up), and environmental factors (e.g., lighting, backgrounds). This is a major drawback for the scalability of these models in VR. This paper makes progress in overcoming these limitations by proposing an end-to-end multi-identity architecture (MIA) trained with specialized augmentation strategies. MIA drives the shape component of the avatar from three cameras in the VR headset (two eyes, one mouth), in untrained subjects, using minimal personalized information (i.e., neutral 3D mesh shape). Similarly, if the PS texture decoder is available, MIA is able to drive the full avatar (shape + texture) robustly outperforming PS models in challenging scenarios. Our key contribution to improve robustness and generalization, is that our method implicitly decouples, in an unsupervised manner, the facial expression from nuisance factors (e.g., headset, environment, facial appearance). We demonstrate the superior performance and robustness of the proposed method versus state-of-the-art PS approaches in a variety of experiments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b5a5c69b-a904-4ff3-b36b-f7755cc42c6eBuilds on7
- Learning an animatable detailed 3D face model from in-the-wild imagesYao Feng, Haiwen Feng, Michael J. Black, Timo BolkartSIGGRAPH 2021 · 662 citations
- The eyes have it: an integrated eye and face model for photorealistic facial animationGabriel Schwartz, Shih-En Wei, Te-Li Wang, Stephen Lombardi et al.SIGGRAPH 2020 · 54 citations
- EgoRenderer: Rendering Human Avatars from Egocentric Camera ImagesTao Hu, Kripasindhu Sarkar, Lingjie Liu, Matthias Zwicker et al.ICCV 2021 · 23 citations
- Unsupervised Learning Facial Parameter Regressor for Action Unit Intensity Estimation via Differentiable RendererXinhui Song, Tianyang Shi, Zunlei Feng, Mingli Song et al.ACM MM 2020 · 6 citations
- High-Fidelity Face Tracking for AR/VR via Deep Lighting AdaptationLele Chen, Chen Cao, Fernando De la Torre, Jason M. Saragih et al.CVPR 2021
Related papers
- Universal Facial Encoding of Codec Avatars from VR HeadsetsShaojie Bai, Te-Li Wang, Chenghui Li, Akshay Venkatesh et al.SIGGRAPH 2024 · 4 citations
- Raise Your Eyebrows Higher: Facilitating Emotional Communication in Social Virtual Reality Through Region-Specific Facial Expression ExaggerationXueyang Wang, Sheng Zhao, Yihe Wang, Howard Ziyu Han et al.CHI 2025 · 11 citations
- Pixel-Aligned Volumetric AvatarsAmit Raj, Michael Zollhöfer, Tomas Simon, Jason M. Saragih et al.CVPR 2021
- Coverage of Facial Expressions and Its Effects on Avatar Embodiment, Self-Identification, and UncanninessPeter Kullmann, Theresa Schell, Timo Menzel, Mario Botsch et al.IEEE VR 2025 · 12 citations
- Pixel Codec AvatarsShugao Ma, Tomas Simon, Jason M. Saragih, Dawei Wang et al.CVPR 2021
