From Inference to Generation: End-to-end Fully Self-supervised Generation of Human Face from Speech
Hyeong-Seok Choi, Changdae Park, Kyogu Lee
2020年份
33被引次数
6顶会引用
摘要
This work seeks the possibility of generating the human face from voice solely based on the audio-visual data without any human-labeled annotations. To this end, we propose a multi-modal learning framework that links the inference stage and generation stage. First, the inference networks are trained to match the speaker identity between the two different modalities. Then the pre-trained inference networks cooperate with the generation network by giving conditional information about the voice.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Neural Analysis and Synthesis: Reconstructing Speech from Self-Supervised RepresentationsHyeong-Seok Choi, Juheon Lee, Wansoo Kim, Jie Lee 等NeurIPS 2021 · 被引用 200 次
- Cross-Modal Perceptionist: Can Face Geometry be Gleaned from Voices?Cho-Ying Wu, Chin-Cheng Hsu, Ulrich NeumannCVPR 2022 · 被引用 16 次
- FaceChain-ImagineID: Freely Crafting High-Fidelity Diverse Talking Faces from Disentangled AudioChao Xu, Yang Liu, Jiazheng Xing, Weida Wang 等CVPR 2024 · 被引用 11 次
- Face-Driven Zero-Shot Voice Conversion with Memory-based Face-Voice AlignmentZhengyan Sheng, Yang Ai, Yan-Nian Chen, Zhen-Hua LingACM MM 2023 · 被引用 5 次
- Posterior Promoted GAN With Distribution Discriminator for Unsupervised Image SynthesisXianchao Zhang, Ziyang Cheng, Xiaotong Zhang, Han LiuCVPR 2021
相关 Paper
- VisualVoice: Audio-Visual Speech Separation With Cross-Modal ConsistencyRuohan Gao, Kristen GraumanCVPR 2021
- Face-based Voice Conversion: Learning the Voice behind a FaceHsiao-Han Lu, Shao-En Weng, Ya-Fan Yen, Hong-Han Shuai 等ACM MM 2021 · 被引用 15 次
- Pose-Controllable Talking Face Generation by Implicitly Modularized Audio-Visual RepresentationHang Zhou, Yasheng Sun, Wayne Wu, Chen Change Loy 等CVPR 2021
- Faces that Speak: Jointly Synthesising Talking Face and Speech from TextYoungjoon Jang, Ji-Hoon Kim, Junseok Ahn, Doyeop Kwak 等CVPR 2024
- Hearing like Seeing: Improving Voice-Face Interactions and Associations via Adversarial Deep Semantic Matching NetworkKai Cheng, Xin Liu, Yiu-ming Cheung, Rui Wang 等ACM MM 2020 · 被引用 19 次
