Real-time 3D neural facial animation from binocular video
Chen Cao, Vasu Agrawal, Fernando De la Torre, Lele Chen, Jason M. Saragih, Tomas Simon, Yaser Sheikh
Abstract
We present a method for performing real-time facial animation of a 3D avatar from binocular video. Existing facial animation methods fail to automatically capture precise and subtle facial motions for driving a photo-realistic 3D avatar "in-the-wild" (i.e., variability in illumination, camera noise). The novelty of our approach lies in a light-weight process for specializing a personalized face model to new environments that enables extremely accurate real-time face tracking anywhere. Our method uses a pre-trained high-fidelity personalized model of the face that we complement with a novel illumination model to account for variations due to lighting and other factors often encountered in-the-wild (e.g., facial hair growth, makeup, skin blemishes). Our approach comprises two steps. First, we solve for our illumination model's parameters by applying analysis-by-synthesis on a short video recording. Using the pairs of model parameters (rigid, non-rigid) and the original images, we learn a regression for real-time inference from the image space to the 3D shape and texture of the avatar. Second, given a new video, we fine-tune the real-time regression model with a few-shot learning strategy to adapt the regression model to the new environment. We demonstrate our system's ability to precisely capture subtle facial motions in unconstrained scenarios, in comparison to competing methods, on a diverse collection of identities, expressions, and real-world environments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 24ffd427-c577-4b7a-ba2b-ebd6832e57ebCited by top-tier papers1
Ask how each one uses itBuilds on2
- Few-Shot Adversarial Learning of Realistic Neural Talking Head ModelsEgor Zakharov, Aliaksandra Shysheya, Egor Burkov, Victor S. LempitskyICCV 2019 · 687 citations
- The eyes have it: an integrated eye and face model for photorealistic facial animationGabriel Schwartz, Shih-En Wei, Te-Li Wang, Stephen Lombardi et al.SIGGRAPH 2020 · 54 citations
Related papers
- Universal Facial Encoding of Codec Avatars from VR HeadsetsShaojie Bai, Te-Li Wang, Chenghui Li, Akshay Venkatesh et al.SIGGRAPH 2024 · 4 citations
- FlashAvatar: High-Fidelity Head Avatar with Efficient Gaussian EmbeddingJun Xiang, Xuan Gao, Yudong Guo, Juyong ZhangCVPR 2024 · 51 citations
- Vid2Avatar-Pro: Authentic Avatar from Videos in the Wild via Universal PriorChen Guo, Junxuan Li, Yash Kant, Yaser Sheikh et al.CVPR 2025
- EyeNeRF: a hybrid representation for photorealistic synthesis, animation and relighting of human eyesGengyan Li, Abhimitra Meka, Franziska Mueller, Marcel C. Bühler et al.SIGGRAPH 2022 · 39 citations
- Real-Time 3D-Aware Portrait Video RelightingZiqi Cai, Kaiwen Jiang, Shu-Yu Chen, Yu-Kun Lai et al.CVPR 2024
