Speech Fusion to Face: Bridging the Gap Between Human's Vocal Characteristics and Facial Imaging
Yeqi Bai, Tao Ma, Lipo Wang, Zhenjie Zhang
摘要
While deep learning technologies are now capable of generating realistic images confusing humans, the research efforts are turning to the synthesis of images for more concrete and application-specific purposes. Facial image generation based on vocal characteristics from speech is one of such important yet challenging tasks. It is the key enabler to influential use cases of image generation, especially for business in public security and entertainment. Existing solutions to the problem of speech2face renders limited image quality and fails to preserve facial similarity due to the lack of quality dataset for training and appropriate integration of vocal features. In this paper, we investigate these key technical challenges and propose Speech Fusion to Face, or SF2F in short, attempting to address the issue of facial image quality and the poor connection between vocal feature domain and modern image generation models. By adopting new strategies on data model and training, we demonstrate dramatic performance boost over state-of-the-art solution, by doubling the recall of individual identity, and lifting the quality score from 15 to 19 based on the mutual information score with VGGFace classifier.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Diffusion Facial Forgery DetectionHarry Cheng, Yangyang Guo, Tianyi Wang, Liqiang Nie 等ACM MM 2024 · 被引用 42 次
- FaceChain-ImagineID: Freely Crafting High-Fidelity Diverse Talking Faces from Disentangled AudioChao Xu, Yang Liu, Jiazheng Xing, Weida Wang 等CVPR 2024 · 被引用 11 次
- Face-Driven Zero-Shot Voice Conversion with Memory-based Face-Voice AlignmentZhengyan Sheng, Yang Ai, Yan-Nian Chen, Zhen-Hua LingACM MM 2023 · 被引用 5 次
- Seeking the Shape of Sound: An Adaptive Framework for Learning Voice-Face AssociationPeisong Wen, Qianqian Xu, Yangbangyan Jiang, Zhiyong Yang 等CVPR 2021
它引用的顶会 Paper1
相关 Paper
- Seeing Your Speech Style: A Novel Zero-Shot Identity-Disentanglement Face-based Voice ConversionYan Rong, Li LiuAAAI 2025 · 被引用 11 次
- VIGFace: Virtual Identity Generation for Privacy-Free Face Recognition DatasetMinsoo Kim, Min-Cheol Sagong, Gi Pyo Nam, Junghyun Cho 等ICCV 2025 · 被引用 1 次
- What Does Your Face Sound Like? 3D Face Shape towards VoiceZhihan Yang, Zhiyong Wu, Ying Shan, Jia JiaAAAI 2023 · 被引用 6 次
- EmoTalk: Speech-Driven Emotional Disentanglement for 3D Face AnimationZiqiao Peng, Haoyu Wu, Zhenbo Song, Hao Xu 等ICCV 2023 · 被引用 192 次
- SynFace: Face Recognition with Synthetic DataHaibo Qiu, Baosheng Yu, Dihong Gong, Zhifeng Li 等ICCV 2021 · 被引用 162 次
