What Does Your Face Sound Like? 3D Face Shape towards Voice
Zhihan Yang, Zhiyong Wu, Ying Shan, Jia Jia
摘要
Face-based speech synthesis provides a practical solution to generate voices from human faces. However, directly using 2D face images leads to the problems of uninterpretability and entanglement. In this paper, to address the issues, we introduce 3D face shape which (1) has an anatomical relationship between voice characteristics, partaking in the "bone conduction" of human timbre production, and ( 2 ) is naturally independent of irrelevant factors by excluding the blending process. We devise a three-stage framework to generate speech from 3D face shapes. Fully considering timbre production in anatomical and acquired terms, our framework incorporates three additional relevant attributes including face texture, facial features, and demographics. Experiments and subjective tests demonstrate our method can generate utterances matching faces well, with good audio quality and voice diversity. We also explore and visualize how the voice changes with the face. Case studies show that our method upgrades the face-voice inference to personalized custom-made voice creating, revealing a promising prospect in virtual human and dubbing applications.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Can I Hear Your Face? Pervasive Attack on Voice Authentication Systems with a Single Face ImageNan Jiang, Bangjie Sun, Terence Sim, Jun HanUSENIX Security 2024 · 被引用 7 次
- Face-Driven Zero-Shot Voice Conversion with Memory-based Face-Voice AlignmentZhengyan Sheng, Yang Ai, Yan-Nian Chen, Zhen-Hua LingACM MM 2023 · 被引用 5 次
它引用的顶会 Paper3
- Learning an animatable detailed 3D face model from in-the-wild imagesYao Feng, Haiwen Feng, Michael J. Black, Timo BolkartSIGGRAPH 2021 · 被引用 662 次
- Single-shot high-quality facial geometry and skin appearance captureJérémy Riviere, Paulo F. U. Gotardo, Derek Bradley, Abhijeet Ghosh 等SIGGRAPH 2020 · 被引用 66 次
- Face-based Voice Conversion: Learning the Voice behind a FaceHsiao-Han Lu, Shao-En Weng, Ya-Fan Yen, Hong-Han Shuai 等ACM MM 2021 · 被引用 15 次
相关 Paper
- Zero-Shot Face-Based Voice Conversion: Bottleneck-Free Speech Disentanglement in the Real-World ScenarioShao-En Weng, Hong-Han Shuai, Wen-Huang ChengAAAI 2023 · 被引用 4 次
- Rethinking Voice-Face Correlation: A Geometry ViewXiang Li, Yandong Wen, Muqiao Yang, Jinglu Wang 等ACM MM 2023 · 被引用 4 次
- MeshTalk: 3D Face Animation from Speech using Cross-Modality DisentanglementAlexander Richard, Michael Zollhöfer, Yandong Wen, Fernando De la Torre 等ICCV 2021 · 被引用 272 次
- Speech-Driven 3D Face Animation with Composite and Regional Facial MovementsHaozhe Wu, Songtao Zhou, Jia Jia, Junliang Xing 等ACM MM 2023 · 被引用 20 次
- Emotional Face-to-SpeechJiaxin Ye, Boyuan Cao, Hongming ShanICML 2025
