AVFace: Towards Detailed Audio-Visual 4D Face Reconstruction
Aggelina Chatziagapi, Dimitris Samaras
Abstract
In this work, we present a multimodal solution to the problem of 4D face reconstruction from monocular videos. 3D face reconstruction from 2D images is an under-constrained problem due to the ambiguity of depth. State-of-the-art methods try to solve this problem by leveraging visual information from a single image or video, whereas 3D mesh animation approaches rely more on audio. However, in most cases (e.g. AR/VR applications), videos include both visual and speech information. We propose AV-Face that incorporates both modalities and accurately re-constructs the 4D facial and lip motion of any speaker, without requiring any 3D ground truth for training. A coarse stage estimates the per-frame parameters of a 3D mor-phable model, followed by a lip refinement, and then a fine stage recovers facial geometric details. Due to the temporal audio and video information captured by transformer-based modules, our method is robust in cases when either modality is insufficient (e.g. face occlusions). Extensive qualitative and quantitative evaluation demonstrates the superiority of our method over the current state-of-the-art.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a7318ec4-ef3c-48b7-8833-cc3348a9daccCited by top-tier papers2
- MemoryTalker: Personalized Speech-Driven 3D Facial Animation via Audio-Guided StylizationHyung Kyu Kim, Sangmin Lee, Hak Gu KimICCV 2025 · 1 citation
- DICE: End-to-end Deformation Capture of Hand-Face Interactions from a Single ImageQingxuan Wu, Zhiyang Dou, Sirui Xu, Soshi Shimada et al.ICLR 2025
Builds on16
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- Implicit Neural Representations with Periodic Activation FunctionsVincent Sitzmann, Julien N. P. Martel, Alexander W. Bergman, David B. Lindell et al.NeurIPS 2020 · 4,008 citations
- Nerfies: Deformable Neural Radiance FieldsKeunhong Park, Utkarsh Sinha, Jonathan T. Barron, Sofien Bouaziz et al.ICCV 2021 · 1,442 citations
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 869 citations
- Learning an animatable detailed 3D face model from in-the-wild imagesYao Feng, Haiwen Feng, Michael J. Black, Timo BolkartSIGGRAPH 2021 · 662 citations
Related papers
- Speech4Mesh: Speech-Assisted Monocular 3D Facial Reconstruction for Speech-Driven 3D Facial AnimationShan He, Haonan He, Shuo Yang, Xiaoyan Wu et al.ICCV 2023 · 13 citations
- Semi-supervised Speech-driven 3D Facial Animation via Cross-modal EncodingPeiji Yang, Huawei Wei, Yicheng Zhong, Zhisheng WangICCV 2023 · 1 citation
- Accurate 3D Face Reconstruction with Facial Component TokensTianke Zhang, Xuangeng Chu, Yunfei Liu, Lijian Lin et al.ICCV 2023 · 38 citations
- V2M4: 4D Mesh Animation Reconstruction from a Single Monocular VideoJianqi Chen, Biao Zhang, Xiangjun Tang, Peter WonkaICCV 2025 · 8 citations
- AV-Flow: Transforming Text to Audio-Visual Human-Like InteractionsAggelina Chatziagapi, Louis-Philippe Morency, Hongyu Gong, Michael Zollhöfer et al.ICCV 2025 · 3 citations
