Lip to Speech Synthesis with Visual Context Attentional GAN
Minsu Kim, Joanna Hong, Yong Man Ro
摘要
In this paper, we propose a novel lip-to-speech generative adversarial network, Visual Context Attentional GAN (VCA-GAN), which can jointly model local and global lip movements during speech synthesis. Specifically, the proposed VCA-GAN synthesizes the speech from local lip visual features by finding a mapping function of viseme-to-phoneme, while global visual context is embedded into the intermediate layers of the generator to clarify the ambiguity in the mapping induced by homophene. To achieve this, a visual context attention module is proposed where it encodes global representations from the local visual features, and provides the desired global visual context corresponding to the given coarse speech representation to the generator through audio-visual attention. In addition to the explicit modelling of local and global visual representations, synchronization learning is introduced as a form of contrastive learning that guides the generator to synthesize a speech in sync with the given input lip movements. Extensive experiments demonstrate that the proposed VCA-GAN outperforms existing state-of-the-art and is able to effectively synthesize the speech from multi-speaker that has been barely handled in the previous works.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Distinguishing Homophenes Using Multi-Head Visual-Audio Memory for Lip ReadingMinsu Kim, Jeong Hun Yeo, Yong Man RoAAAI 2022 · 被引用 86 次
- DiffV2S: Diffusion-based Video-to-Speech Synthesis with Vision-guided Speaker EmbeddingJeongsoo Choi, Joanna Hong, Yong Man RoICCV 2023 · 被引用 34 次
- Lip Reading for Low-resource Languages by Learning and Combining General Speech Knowledge and Language-specific KnowledgeMinsu Kim, Jeong Hun Yeo, Jeongsoo Choi, Yong Man RoICCV 2023 · 被引用 31 次
- LipVoicer: Generating Speech from Silent Videos Guided by Lip ReadingYochai Yemini, Aviv Shamsian, Lior Bracha, Sharon Gannot 等ICLR 2024 · 被引用 29 次
- Towards Accurate Lip-to-Speech Synthesis in-the-WildSindhu B. Hegde, Rudrabha Mukhopadhyay, C. V. Jawahar, Vinay P. NamboodiriACM MM 2023 · 被引用 9 次
它引用的顶会 Paper6
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 被引用 2,890 次
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 被引用 869 次
- Multi-modality Associative Bridging through Memory: Speech Sound Recollected from Face VideoMinsu Kim, Joanna Hong, Se Jin Park, Yong Man RoICCV 2021 · 被引用 48 次
- Discriminative Multi-Modality Speech RecognitionBo Xu, Cheng Lu, Yandong Guo, Jacob WangCVPR 2020
相关 Paper
- FACIAL: Synthesizing Dynamic Talking Face with Implicit Attribute LearningChenxu Zhang, Yifan Zhao, Yifei Huang, Ming Zeng 等ICCV 2021 · 被引用 149 次
- SyncTalkFace: Talking Face Generation with Precise Lip-Syncing via Audio-Lip MemorySe Jin Park, Minsu Kim, Joanna Hong, Jeongsoo Choi 等AAAI 2022 · 被引用 110 次
- Lip-to-Speech Synthesis for Arbitrary Speakers in the WildSindhu B. Hegde, K. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri 等ACM MM 2022 · 被引用 15 次
- Seeing What You Said: Talking Face Generation Guided by a Lip Reading ExpertJiadong Wang, Xinyuan Qian, Malu Zhang, Robby T. Tan 等CVPR 2023
- Speech2Lip: High-fidelity Speech to Lip Generation by Learning from a Short VideoXiuzhe Wu, Pengfei Hu, Yang Wu, Xiaoyang Lyu 等ICCV 2023 · 被引用 18 次
