From Inference to Generation: End-to-end Fully Self-supervised Generation of Human Face from Speech
Hyeong-Seok Choi, Changdae Park, Kyogu Lee
2020Year
33Citations
6Top-tier citations
Abstract
This work seeks the possibility of generating the human face from voice solely based on the audio-visual data without any human-labeled annotations. To this end, we propose a multi-modal learning framework that links the inference stage and generation stage. First, the inference networks are trained to match the speaker identity between the two different modalities. Then the pre-trained inference networks cooperate with the generation network by giving conditional information about the voice.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Neural Analysis and Synthesis: Reconstructing Speech from Self-Supervised RepresentationsHyeong-Seok Choi, Juheon Lee, Wansoo Kim, Jie Lee et al.NeurIPS 2021 · 200 citations
- Cross-Modal Perceptionist: Can Face Geometry be Gleaned from Voices?Cho-Ying Wu, Chin-Cheng Hsu, Ulrich NeumannCVPR 2022 · 16 citations
- FaceChain-ImagineID: Freely Crafting High-Fidelity Diverse Talking Faces from Disentangled AudioChao Xu, Yang Liu, Jiazheng Xing, Weida Wang et al.CVPR 2024 · 11 citations
- Face-Driven Zero-Shot Voice Conversion with Memory-based Face-Voice AlignmentZhengyan Sheng, Yang Ai, Yan-Nian Chen, Zhen-Hua LingACM MM 2023 · 5 citations
- Posterior Promoted GAN With Distribution Discriminator for Unsupervised Image SynthesisXianchao Zhang, Ziyang Cheng, Xiaotong Zhang, Han LiuCVPR 2021
Related papers
- VisualVoice: Audio-Visual Speech Separation With Cross-Modal ConsistencyRuohan Gao, Kristen GraumanCVPR 2021
- Face-based Voice Conversion: Learning the Voice behind a FaceHsiao-Han Lu, Shao-En Weng, Ya-Fan Yen, Hong-Han Shuai et al.ACM MM 2021 · 15 citations
- Pose-Controllable Talking Face Generation by Implicitly Modularized Audio-Visual RepresentationHang Zhou, Yasheng Sun, Wayne Wu, Chen Change Loy et al.CVPR 2021
- Faces that Speak: Jointly Synthesising Talking Face and Speech from TextYoungjoon Jang, Ji-Hoon Kim, Junseok Ahn, Doyeop Kwak et al.CVPR 2024
- Hearing like Seeing: Improving Voice-Face Interactions and Associations via Adversarial Deep Semantic Matching NetworkKai Cheng, Xin Liu, Yiu-ming Cheung, Rui Wang et al.ACM MM 2020 · 19 citations
