Learning Individual Speaking Styles for Accurate Lip to Speech Synthesis
K. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. Jawahar
Abstract
Humans involuntarily tend to infer parts of the conversation from lip movements when the speech is absent or corrupted by external noise. In this work, we explore the task of lip to speech synthesis, i.e., learning to generate natural speech given only the lip movements of a speaker. Acknowledging the importance of contextual and speakerspecific cues for accurate lip-reading, we take a different path from existing works. We focus on learning accurate lip sequences to speech mappings for individual speakers in unconstrained, large vocabulary settings. To this end, we collect and release a large-scale benchmark dataset, the first of its kind, specifically to train and evaluate the singlespeaker lip to speech task in natural settings. We propose a novel approach with key design choices to achieve accurate, natural lip to speech synthesis in such unconstrained scenarios for the first time. Extensive evaluation using quantitative, qualitative metrics and human evaluation shows that our method is four times more intelligible than previous works in this space.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 29ed8442-39f6-44f1-9891-502d8dbedcf9Cited by top-tier papers29
- V2A-Mapper: A Lightweight Solution for Vision-to-Audio Generation by Connecting Foundation ModelsHeng Wang, Jianbo Ma, Santiago Pascual, Richard Cartwright et al.AAAI 2024 · 84 citations
- Lip to Speech Synthesis with Visual Context Attentional GANMinsu Kim, Joanna Hong, Yong Man RoNeurIPS 2021 · 76 citations
- Neural Dubber: Dubbing for Videos According to ScriptsChenxu Hu, Qiao Tian, Tingle Li, Yuping Wang et al.NeurIPS 2021 · 62 citations
- Multi-modality Associative Bridging through Memory: Speech Sound Recollected from Face VideoMinsu Kim, Joanna Hong, Se Jin Park, Yong Man RoICCV 2021 · 48 citations
- LipLearner: Customizable Silent Speech Interactions on Mobile DevicesZixiong Su, Shitao Fang, Jun RekimotoCHI 2023 · 37 citations
Builds on1
Related papers
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 869 citations
- Lip-to-Speech Synthesis for Arbitrary Speakers in the WildSindhu B. Hegde, K. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri et al.ACM MM 2022 · 15 citations
- Towards Accurate Lip-to-Speech Synthesis in-the-WildSindhu B. Hegde, Rudrabha Mukhopadhyay, C. V. Jawahar, Vinay P. NamboodiriACM MM 2023 · 9 citations
- Let There Be Sound: Reconstructing High Quality Speech from Silent VideosJi-Hoon Kim, Jaehun Kim, Joon Son ChungAAAI 2024 · 14 citations
- Speech2Lip: High-fidelity Speech to Lip Generation by Learning from a Short VideoXiuzhe Wu, Pengfei Hu, Yang Wu, Xiaoyang Lyu et al.ICCV 2023 · 18 citations
