SyncTalklip: Highly Synchronized Lip-Readable Speaker Generation with Multi-Task Learning
Xiaoda Yang, Xize Cheng, Dongjie Fu, Minghui Fang, Jialong Zuo, Shengpeng Ji, Zhou Zhao, Tao Jin
摘要
Talking Face Generation (TFG) reconstructs facial motions concerning lips given speech input, which aims to generate highquality, synchronized, and lip-readable videos. Previous efforts have achieved success in generating quality and synchronization, and recently, there has been an increasing focus on the importance of intelligibility. Despite these efforts, there remains a challenge in achieving a balance among quality, synchronization, and intelligibility, often resulting in trade-offs that compromise one aspect in favor of another. In light of this, we propose SyncTalklip, a novel dual-tower framework designed to overcome the challenges of synchronization while improving lip-reading performance. To enhance the performance of SyncTalklip in both synchronization and intelligibility, we design AV-SyncNet, a pre-trained multi-task model, aiming to achieve a dual-focus on synchronization and intelligibility. Moreover, we propose a novel cross-modal contrastive learning bringing audio and video closer to enhance synchronization. Experimental results demonstrate that SyncTalklip achieves state-of-the-art performance in quality, intelligibility, and synchronization. Furthermore, extensive experiments have demonstrated our model's generalizability across domains. The code and demo is available at https://sync-talklip.github.io.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper6
- Low-rank Prompt Interaction for Continual Vision-Language RetrievalWeicai Yan, Ye Wang, Wang Lin, Zirun Guo 等ACM MM 2024 · 被引用 8 次
- AudioVSR: Enhancing Video Speech Recognition with Audio DataXiaoda Yang, Xize Cheng, Jiaqi Duan, Hongshun Qiu 等EMNLP 2024 · 被引用 3 次
- SpatialLogic-Bench: A Diagnostic Benchmark for Task-Oriented Spatiotemporal ReasoningXiaoda Yang, Shenzhou Gao, Can Wang, Jiahe Zhang 等AAAI 2026
- PACHAT: Persona-Aware Speech Assistant for Multi-party DialogueDongjie Fu, Xize Cheng, Linjun Li, Xiaoda Yang 等EMNLP 2025
- VoxDialogue: Can Spoken Dialogue Systems Understand Information Beyond Words?Xize Cheng, Ruofan Hu, Xiaoda Yang, Jingyu Lu 等ICLR 2025
相关 Paper
- Seeing What You Said: Talking Face Generation Guided by a Lip Reading ExpertJiadong Wang, Xinyuan Qian, Malu Zhang, Robby T. Tan 等CVPR 2023
- DualLip: A System for Joint Lip Reading and GenerationWeicong Chen, Xu Tan, Yingce Xia, Tao Qin 等ACM MM 2020 · 被引用 25 次
- SelfTalk: A Self-Supervised Commutative Training Diagram to Comprehend 3D Talking FacesZiqiao Peng, Yihao Luo, Yue Shi, Hao Xu 等ACM MM 2023 · 被引用 56 次
- Speech2Lip: High-fidelity Speech to Lip Generation by Learning from a Short VideoXiuzhe Wu, Pengfei Hu, Yang Wu, Xiaoyang Lyu 等ICCV 2023 · 被引用 18 次
- SyncTalk: The Devil is in the Synchronization for Talking Head SynthesisZiqiao Peng, Wentao Hu, Yue Shi, Xiangyu Zhu 等CVPR 2024 · 被引用 65 次
