Uni-Dubbing: Zero-Shot Speech Synthesis from Visual Articulation
Songju Lei, Xize Cheng, Mengjiao Lyu, Jianqiao Hu, Jintao Tan, Runlin Liu, Lingyu Xiong, Tao Jin, Xiandong Li, Zhou Zhao
摘要
In the field of speech synthesis, there is a growing emphasis on employing multimodal speech to enhance robustness. A key challenge in this area is the scarcity of datasets that pair audio with corresponding video. We employ a methodology that incorporates modality alignment during the pre-training phase on multimodal datasets, uniquely facilitating Zero-Shot generalization through the process of freezing the video modality feature extraction component and the encoder module within the pretrained weights, thereby enabling effective cross-modal and cross-lingual transfer. We have named this method 'Uni-Dubbing'. Our method finely tunes with both multimodal and single-modality audio data. In multimodal scenarios, it achieves a reduced word error rate (WER) of 31.73%, surpassing the previous best of 33.9%. It also excels in metrics like tone quality and synchronization. With single-modality audio, it achieves a WER of 36.08%, demonstrating adaptability to limited data. Its domain generalization capabilities are proven across various language tasks in video translation and audio generation. Trained on 433 hours of audio data, it surpasses techniques using 200 hours of audiovisual data. The code and demo are available at https://diracer.github.io/unidubbing .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio GenerationYan Rong, Jinting Wang, Guangzhi Lei, Shan Yang 等ACM MM 2025 · 被引用 1 次
- UniCUE: Unified Recognition and Generation Framework for Chinese Cued Speech Video-to-Speech GenerationJinting Wang, Shan Yang, Chenxing Li, Dong Yu 等AAAI 2026
- VoxDialogue: Can Spoken Dialogue Systems Understand Information Beyond Words?Xize Cheng, Ruofan Hu, Xiaoda Yang, Jingyu Lu 等ICLR 2025
- SLD-L2S: Hierarchical Subspace Latent Diffusion for High-Fidelity Lip to Speech SynthesisYifan Liang, Andong Li, Kang Yang, Guochen Yu 等AAAI 2026
- From Faces to Voices: Learning Hierarchical Representations for High-quality Video-to-SpeechJi-Hoon Kim, Jeongsoo Choi, Jaehun Kim, Chaeyoung Jung 等CVPR 2025
它引用的顶会 Paper7
- Translatotron 2: High-quality direct speech-to-speech translation with voice preservationYe Jia, Michelle Tadmor Ramanovich, Tal Remez, Roi PomerantzICML 2022 · 被引用 107 次
- Lip to Speech Synthesis with Visual Context Attentional GANMinsu Kim, Joanna Hong, Yong Man RoNeurIPS 2021 · 被引用 76 次
- UWSpeech: Speech to Speech Translation for Unwritten LanguagesChen Zhang, Xu Tan, Yi Ren, Tao Qin 等AAAI 2021 · 被引用 69 次
- LipVoicer: Generating Speech from Silent Videos Guided by Lip ReadingYochai Yemini, Aviv Shamsian, Lior Bracha, Sharon Gannot 等ICLR 2024 · 被引用 29 次
- Rethinking Missing Modality Learning from a Decoding PerspectiveTao Jin, Xize Cheng, Linjun Li, Wang Lin 等ACM MM 2023 · 被引用 10 次
相关 Paper
- AudioVSR: Enhancing Video Speech Recognition with Audio DataXiaoda Yang, Xize Cheng, Jiaqi Duan, Hongshun Qiu 等EMNLP 2024 · 被引用 3 次
- An eye for an ear: zero-shot audio description leveraging an image captioner with audio-visual token distribution matchingHugo Malard, Michel Olvera, Stéphane Lathuilière, Slim EssidNeurIPS 2024 · 被引用 3 次
- DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio SynthesisWenjie Tian, Xinfa Zhu, Haohe Liu, Zhixian Zhao 等ACM MM 2025
- From Speaker to Dubber: Movie Dubbing with Prosody and Duration Consistency LearningZhedong Zhang, Liang Li, Gaoxiang Cong, Haibing Yin 等ACM MM 2024 · 被引用 36 次
- UniSpeech: Unified Speech Representation Learning with Labeled and Unlabeled DataChengyi Wang, Yu Wu, Yao Qian, Ken'ichi Kumatani 等ICML 2021 · 被引用 140 次
