SyncTalkFace: Talking Face Generation with Precise Lip-Syncing via Audio-Lip Memory
Se Jin Park, Minsu Kim, Joanna Hong, Jeongsoo Choi, Yong Man Ro
Abstract
The challenge of talking face generation from speech lies in aligning two different modal information, audio and video, such that the mouth region corresponds to input audio. Previous methods either exploit audio-visual representation learning or leverage intermediate structural information such as landmarks and 3D models. However, they struggle to synthesize fine details of the lips varying at the phoneme level as they do not sufficiently provide visual information of the lips at the video synthesis step. To overcome this limitation, our work proposes Audio-Lip Memory that brings in visual information of the mouth region corresponding to input audio and enforces fine-grained audio-visual coherence. It stores lip motion features from sequential ground truth images in the value memory and aligns them with corresponding audio features so that they can be retrieved using audio input at inference time. Therefore, using the retrieved lip motion features as visual hints, it can easily correlate audio with visual dynamics in the synthesis step. By analyzing the memory, we demonstrate that unique lip features are stored in each memory slot at the phoneme level, capturing subtle lip motion based on memory addressing. In addition, we introduce visual-visual synchronization loss which can enhance lip-syncing performance when used along with audio-visual synchronization loss in our model. Extensive experiments are performed to verify that our method generates high-quality video with mouth shapes that best align with the input audio, outperforming previous state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 26f69515-f792-4ad4-8a47-b87ff939b8d9Cited by top-tier papers25
- DINet: Deformation Inpainting Network for Realistic Face Visually Dubbing on High Resolution VideoZhimeng Zhang, Zhipeng Hu, Wenjin Deng, Changjie Fan et al.AAAI 2023 · 106 citations
- EMMN: Emotional Motion Memory Network for Audio-driven Emotional Talking Face GenerationShuai Tan, Bin Ji, Ye PanICCV 2023 · 63 citations
- SelfTalk: A Self-Supervised Commutative Training Diagram to Comprehend 3D Talking FacesZiqiao Peng, Yihao Luo, Yue Shi, Hao Xu et al.ACM MM 2023 · 56 citations
- MODA: Mapping-Once Audio-driven Portrait Animation with Dual AttentionsYunfei Liu, Lijian Lin, Fei Yu, Changyin Zhou et al.ICCV 2023 · 40 citations
- Say Anything with Any StyleShuai Tan, Bin Ji, Yu Ding, Ye PanAAAI 2024 · 30 citations
Builds on8
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 869 citations
- Robust Small-scale Pedestrian Detection with Cued Recall via Memory LearningJung Uk Kim, Sungjune Park, Yong Man RoICCV 2021 · 61 citations
- Multi-modality Associative Bridging through Memory: Speech Sound Recollected from Face VideoMinsu Kim, Joanna Hong, Se Jin Park, Yong Man RoICCV 2021 · 48 citations
- One-Shot Free-View Neural Talking-Head Synthesis for Video ConferencingTing-Chun Wang, Arun Mallya, Ming-Yu LiuCVPR 2021
- Flow-Guided One-Shot Talking Face Generation With a High-Resolution Audio-Visual DatasetZhimeng Zhang, Lincheng Li, Yu Ding, Changjie FanCVPR 2021
Related papers
- Seeing What You Said: Talking Face Generation Guided by a Lip Reading ExpertJiadong Wang, Xinyuan Qian, Malu Zhang, Robby T. Tan et al.CVPR 2023
- Identity-Preserving Talking Face Generation with Landmark and Appearance PriorsWeizhi Zhong, Chaowei Fang, Yinqi Cai, Pengxu Wei et al.CVPR 2023
- Distinguishing Homophenes Using Multi-Head Visual-Audio Memory for Lip ReadingMinsu Kim, Jeong Hun Yeo, Yong Man RoAAAI 2022 · 86 citations
- Lip to Speech Synthesis with Visual Context Attentional GANMinsu Kim, Joanna Hong, Yong Man RoNeurIPS 2021 · 76 citations
- SyncTalklip: Highly Synchronized Lip-Readable Speaker Generation with Multi-Task LearningXiaoda Yang, Xize Cheng, Dongjie Fu, Minghui Fang et al.ACM MM 2024 · 4 citations
