Lip-to-Speech Synthesis for Arbitrary Speakers in the Wild
Sindhu B. Hegde, K. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. Jawahar
摘要
In this work, we address the problem of generating speech from silent lip videos for any speaker in the wild. In stark contrast to previous works, our method (i) is not restricted to a fixed number of speakers, (ii) does not explicitly impose constraints on the domain or the vocabulary and (iii) deals with videos that are recorded in the wild as opposed to within laboratory settings. The task presents a host of challenges, with the key one being that many features of the desired target speech, like voice, pitch and linguistic content, cannot be entirely inferred from the silent face video. In order to handle these stochastic variations, we propose a new VAE-GAN architecture that learns to associate the lip and speech sequences amidst the variations. With the help of multiple powerful discriminators that guide the training process, our generator learns to synthesize speech sequences in any voice for the lip movements of any person. Extensive experiments on multiple datasets show that we outperform all baselines by a large margin. Further, our network can be fine-tuned on videos of specific identities to achieve a performance comparable to single-speaker models that are trained on more data. We conduct numerous ablation studies to analyze the effect of different modules of our architecture. We also provide a demo video that demonstrates several qualitative results along with the code and trained models on our website http://cvit.iiit.ac.in/research/projects/cvit-projects/lip-to-speech-synthesis.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- UMMAFormer: A Universal Multimodal-adaptive Transformer Framework for Temporal Forgery LocalizationRui Zhang, Hongxia Wang, Mingshan Du, Hanqing Liu 等ACM MM 2023 · 被引用 42 次
- Towards Accurate Lip-to-Speech Synthesis in-the-WildSindhu B. Hegde, Rudrabha Mukhopadhyay, C. V. Jawahar, Vinay P. NamboodiriACM MM 2023 · 被引用 9 次
- Learning to Dub Movies via Hierarchical Prosody ModelsGaoxiang Cong, Liang Li, Yuankai Qi, Zheng-Jun Zha 等CVPR 2023
- From Faces to Voices: Learning Hierarchical Representations for High-quality Video-to-SpeechJi-Hoon Kim, Jeongsoo Choi, Jaehun Kim, Chaeyoung Jung 等CVPR 2025
它引用的顶会 Paper5
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 被引用 869 次
- High Fidelity Speech Synthesis with Adversarial NetworksMikolaj Binkowski, Jeff Donahue, Sander Dieleman, Aidan Clark 等ICLR 2020 · 被引用 263 次
- Sub-word Level Lip Reading With Visual AttentionK. R. Prajwal, Triantafyllos Afouras, Andrew ZissermanCVPR 2022 · 被引用 104 次
- Learning Individual Speaking Styles for Accurate Lip to Speech SynthesisK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharCVPR 2020
相关 Paper
- Lip to Speech Synthesis with Visual Context Attentional GANMinsu Kim, Joanna Hong, Yong Man RoNeurIPS 2021 · 被引用 76 次
- FastLTS: Non-Autoregressive End-to-End Unconstrained Lip-to-Speech SynthesisYongqi Wang, Zhou ZhaoACM MM 2022 · 被引用 9 次
- Let There Be Sound: Reconstructing High Quality Speech from Silent VideosJi-Hoon Kim, Jaehun Kim, Joon Son ChungAAAI 2024 · 被引用 14 次
- VisualVoice: Audio-Visual Speech Separation With Cross-Modal ConsistencyRuohan Gao, Kristen GraumanCVPR 2021
- Flow-Based Unconstrained Lip to Speech GenerationJinzheng He, Zhou Zhao, Yi Ren, Jinglin Liu 等AAAI 2022 · 被引用 21 次
