What Does Your Face Sound Like? 3D Face Shape towards Voice
Zhihan Yang, Zhiyong Wu, Ying Shan, Jia Jia
Abstract
Face-based speech synthesis provides a practical solution to generate voices from human faces. However, directly using 2D face images leads to the problems of uninterpretability and entanglement. In this paper, to address the issues, we introduce 3D face shape which (1) has an anatomical relationship between voice characteristics, partaking in the "bone conduction" of human timbre production, and ( 2 ) is naturally independent of irrelevant factors by excluding the blending process. We devise a three-stage framework to generate speech from 3D face shapes. Fully considering timbre production in anatomical and acquired terms, our framework incorporates three additional relevant attributes including face texture, facial features, and demographics. Experiments and subjective tests demonstrate our method can generate utterances matching faces well, with good audio quality and voice diversity. We also explore and visualize how the voice changes with the face. Case studies show that our method upgrades the face-voice inference to personalized custom-made voice creating, revealing a promising prospect in virtual human and dubbing applications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2b010984-4ae9-47d9-867b-b20c3beaf6d7Cited by top-tier papers2
- Can I Hear Your Face? Pervasive Attack on Voice Authentication Systems with a Single Face ImageNan Jiang, Bangjie Sun, Terence Sim, Jun HanUSENIX Security 2024 · 7 citations
- Face-Driven Zero-Shot Voice Conversion with Memory-based Face-Voice AlignmentZhengyan Sheng, Yang Ai, Yan-Nian Chen, Zhen-Hua LingACM MM 2023 · 5 citations
Builds on3
- Learning an animatable detailed 3D face model from in-the-wild imagesYao Feng, Haiwen Feng, Michael J. Black, Timo BolkartSIGGRAPH 2021 · 662 citations
- Single-shot high-quality facial geometry and skin appearance captureJérémy Riviere, Paulo F. U. Gotardo, Derek Bradley, Abhijeet Ghosh et al.SIGGRAPH 2020 · 66 citations
- Face-based Voice Conversion: Learning the Voice behind a FaceHsiao-Han Lu, Shao-En Weng, Ya-Fan Yen, Hong-Han Shuai et al.ACM MM 2021 · 15 citations
Related papers
- Zero-Shot Face-Based Voice Conversion: Bottleneck-Free Speech Disentanglement in the Real-World ScenarioShao-En Weng, Hong-Han Shuai, Wen-Huang ChengAAAI 2023 · 4 citations
- Rethinking Voice-Face Correlation: A Geometry ViewXiang Li, Yandong Wen, Muqiao Yang, Jinglu Wang et al.ACM MM 2023 · 4 citations
- MeshTalk: 3D Face Animation from Speech using Cross-Modality DisentanglementAlexander Richard, Michael Zollhöfer, Yandong Wen, Fernando De la Torre et al.ICCV 2021 · 272 citations
- Speech-Driven 3D Face Animation with Composite and Regional Facial MovementsHaozhe Wu, Songtao Zhou, Jia Jia, Junliang Xing et al.ACM MM 2023 · 20 citations
- Emotional Face-to-SpeechJiaxin Ye, Boyuan Cao, Hongming ShanICML 2025
