Universal Facial Encoding of Codec Avatars from VR Headsets
Shaojie Bai, Te-Li Wang, Chenghui Li, Akshay Venkatesh, Tomas Simon, Chen Cao, Gabriel Schwartz, Jason M. Saragih, Yaser Sheikh, Shih-En Wei
Abstract
Faithful real-time facial animation is essential for avatar-mediated telepresence in Virtual Reality (VR). To emulate authentic communication, avatar animation needs to be efficient and accurate: able to capture both extreme and subtle expressions within a few milliseconds to sustain the rhythm of natural conversations. The oblique and incomplete views of the face, variability in the donning of headsets, and illumination variation due to the environment are some of the unique challenges in generalization to unseen faces. In this paper, we present a method that can animate a photorealistic avatar in realtime from head-mounted cameras (HMCs) on a consumer VR headset. We present a self-supervised learning approach, based on a cross-view reconstruction objective, that enables generalization to unseen users. We present a lightweight expression calibration mechanism that increases accuracy with minimal additional cost to run-time efficiency. We present an improved parameterization for precise ground-truth generation that provides robustness to environmental variation. The resulting system produces accurate facial animation for unseen users wearing VR headsets in realtime. We compare our approach to prior face-encoding methods demonstrating significant improvements in both quantitative metrics and qualitative results.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ef458d53-fbcc-4fa9-84eb-81ec6fc32d7eCited by top-tier papers5
- Enhancing Social Experiences in Immersive Virtual Reality with Artificial Facial MimicryAlessandro Visconti, Davide Calandra, Federica Giorgione, Fabrizio LambertiIEEE VR 2025 · 11 citations
- DeX-Portrait: Disentangled and Expressive Portrait Animation via Explicit and Latent Motion RepresentationsYuxiang Shi, Zhe Li, Yanwen Wang, Hao Zhu et al.CVPR 2026 · 3 citations
- OFERA: Blendshape-Driven 3D Gaussian Control for Occluded Facial Expression to Realistic Avatars in VRSeokhwan Yang, Boram Yoon, Seoyoung Kang, Hail Song et al.IEEE VR 2026 · 1 citation
- EgoRelight: Egocentric Human Capture and Illumination Recovery for Relightable and Photoreal Avatar RenderingJianchun Chen, Yinda Zhang, Rohit Pandey, Thabo Beeler et al.SIGGRAPH 2026
- REWIND: Real-Time Egocentric Whole-Body Motion Diffusion with Exemplar-Based Identity ConditioningJihyun Lee, Weipeng Xu, Alexander Richard, Shih-En Wei et al.CVPR 2025
Builds on19
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- Barlow Twins: Self-Supervised Learning via Redundancy ReductionJure Zbontar, Li Jing, Ishan Misra, Yann LeCun et al.ICML 2021 · 2,942 citations
Related papers
- The eyes have it: an integrated eye and face model for photorealistic facial animationGabriel Schwartz, Shih-En Wei, Te-Li Wang, Stephen Lombardi et al.SIGGRAPH 2020 · 54 citations
- Real-time 3D neural facial animation from binocular videoChen Cao, Vasu Agrawal, Fernando De la Torre, Lele Chen et al.SIGGRAPH 2021 · 29 citations
- Pixel-Aligned Volumetric AvatarsAmit Raj, Michael Zollhöfer, Tomas Simon, Jason M. Saragih et al.CVPR 2021
- VOODOO 3D: Volumetric Portrait Disentanglement for One-Shot 3D Head ReenactmentPhong Tran, Egor Zakharov, Long-Nhat Ho, Anh Tuan Tran et al.CVPR 2024 · 15 citations
- Deep relightable appearance models for animatable facesSai Bi, Stephen Lombardi, Shunsuke Saito, Tomas Simon et al.SIGGRAPH 2021 · 75 citations
