Faces that Speak: Jointly Synthesising Talking Face and Speech from Text
Youngjoon Jang, Ji-Hoon Kim, Junseok Ahn, Doyeop Kwak, Hongsun Yang, Yooncheol Ju, Ilhwan Kim, Byeong-Yeol Kim, Joon Son Chung
2024年份
9顶会引用
摘要
Motion condition Speaker condition Figure 1. Our framework integrates Talking Face Generation (TFG) and Text-to-Speech (TTS) systems, generating synchronised natural speech and a talking face video from a single portrait and text input. Our model is capable of variational motion generation by conditioning the TFG model with the intermediate representations of the TTS model. The speech is conditioned using the identity features extracted in the TFG model to align with the input identity.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- OmniTalker: One-shot Real-time Text-Driven Talking Audio-Video Generation With Multimodal Style MimickingZhongjian Wang, Peng Zhang, Jinwei Qi, Yuan Wang 等NeurIPS 2025 · 被引用 12 次
- ReMask-Animate: Refined Character Image Animation Using Mask-Guided AdaptersXunzhi Xiang, Haiwei Xue, Zonghong Dai, Di Wang 等AAAI 2025 · 被引用 4 次
- AV-Flow: Transforming Text to Audio-Visual Human-Like InteractionsAggelina Chatziagapi, Louis-Philippe Morency, Hongyu Gong, Michael Zollhöfer 等ICCV 2025 · 被引用 3 次
- Hierarchical Codec Diffusion for Video-to-Speech GenerationJiaxin Ye, Gaoxiang Cong, Chenhui Wang, Xin-Cheng Wen 等CVPR 2026 · 被引用 3 次
- FlowDubber: Movie Dubbing with LLM-based Semantic-aware Learning and Flow Matching based Voice EnhancingGaoxiang Cong, Liang Li, Jiadong Pan, Zhedong Zhang 等ACM MM 2025 · 被引用 2 次
它引用的顶会 Paper24
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 被引用 2,890 次
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 被引用 869 次
- Grad-TTS: A Diffusion Probabilistic Model for Text-to-SpeechVadim Popov, Ivan Vovk, Vladimir Gogoryan, Tasnima Sadekova 等ICML 2021 · 被引用 715 次
- Glow-TTS: A Generative Flow for Text-to-Speech via Monotonic Alignment SearchJaehyeon Kim, Sungwon Kim, Jungil Kong, Sungroh YoonNeurIPS 2020 · 被引用 663 次
相关 Paper
- Pose-Controllable Talking Face Generation by Implicitly Modularized Audio-Visual RepresentationHang Zhou, Yasheng Sun, Wayne Wu, Chen Change Loy 等CVPR 2021
- More than Words: In-the-Wild Visually-Driven Prosody for Text-to-SpeechMichael Hassid, Michelle Tadmor Ramanovich, Brendan Shillingford, Miaosen Wang 等CVPR 2022 · 被引用 13 次
- That's What I Said: Fully-Controllable Talking Face GenerationYoungjoon Jang, Kyeongha Rho, Jong-Bin Woo, Hyeongkeun Lee 等ACM MM 2023 · 被引用 7 次
- Write-a-speaker: Text-based Emotional and Rhythmic Talking-head GenerationLincheng Li, Suzhen Wang, Zhimeng Zhang, Yu Ding 等AAAI 2021 · 被引用 88 次
- SyncTalklip: Highly Synchronized Lip-Readable Speaker Generation with Multi-Task LearningXiaoda Yang, Xize Cheng, Dongjie Fu, Minghui Fang 等ACM MM 2024 · 被引用 4 次
