Uni-Dubbing: Zero-Shot Speech Synthesis from Visual Articulation
Songju Lei, Xize Cheng, Mengjiao Lyu, Jianqiao Hu, Jintao Tan, Runlin Liu, Lingyu Xiong, Tao Jin, Xiandong Li, Zhou Zhao
Abstract
In the field of speech synthesis, there is a growing emphasis on employing multimodal speech to enhance robustness. A key challenge in this area is the scarcity of datasets that pair audio with corresponding video. We employ a methodology that incorporates modality alignment during the pre-training phase on multimodal datasets, uniquely facilitating Zero-Shot generalization through the process of freezing the video modality feature extraction component and the encoder module within the pretrained weights, thereby enabling effective cross-modal and cross-lingual transfer. We have named this method 'Uni-Dubbing'. Our method finely tunes with both multimodal and single-modality audio data. In multimodal scenarios, it achieves a reduced word error rate (WER) of 31.73%, surpassing the previous best of 33.9%. It also excels in metrics like tone quality and synchronization. With single-modality audio, it achieves a WER of 36.08%, demonstrating adaptability to limited data. Its domain generalization capabilities are proven across various language tasks in video translation and audio generation. Trained on 433 hours of audio data, it surpasses techniques using 200 hours of audiovisual data. The code and demo are available at https://diracer.github.io/unidubbing .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aa53ce8a-6018-462e-a2a3-e8f9621a4356Cited by top-tier papers6
- AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio GenerationYan Rong, Jinting Wang, Guangzhi Lei, Shan Yang et al.ACM MM 2025 · 1 citation
- UniCUE: Unified Recognition and Generation Framework for Chinese Cued Speech Video-to-Speech GenerationJinting Wang, Shan Yang, Chenxing Li, Dong Yu et al.AAAI 2026
- VoxDialogue: Can Spoken Dialogue Systems Understand Information Beyond Words?Xize Cheng, Ruofan Hu, Xiaoda Yang, Jingyu Lu et al.ICLR 2025
- SLD-L2S: Hierarchical Subspace Latent Diffusion for High-Fidelity Lip to Speech SynthesisYifan Liang, Andong Li, Kang Yang, Guochen Yu et al.AAAI 2026
- From Faces to Voices: Learning Hierarchical Representations for High-quality Video-to-SpeechJi-Hoon Kim, Jeongsoo Choi, Jaehun Kim, Chaeyoung Jung et al.CVPR 2025
Builds on7
- Translatotron 2: High-quality direct speech-to-speech translation with voice preservationYe Jia, Michelle Tadmor Ramanovich, Tal Remez, Roi PomerantzICML 2022 · 107 citations
- Lip to Speech Synthesis with Visual Context Attentional GANMinsu Kim, Joanna Hong, Yong Man RoNeurIPS 2021 · 76 citations
- UWSpeech: Speech to Speech Translation for Unwritten LanguagesChen Zhang, Xu Tan, Yi Ren, Tao Qin et al.AAAI 2021 · 69 citations
- LipVoicer: Generating Speech from Silent Videos Guided by Lip ReadingYochai Yemini, Aviv Shamsian, Lior Bracha, Sharon Gannot et al.ICLR 2024 · 29 citations
- Rethinking Missing Modality Learning from a Decoding PerspectiveTao Jin, Xize Cheng, Linjun Li, Wang Lin et al.ACM MM 2023 · 10 citations
Related papers
- AudioVSR: Enhancing Video Speech Recognition with Audio DataXiaoda Yang, Xize Cheng, Jiaqi Duan, Hongshun Qiu et al.EMNLP 2024 · 3 citations
- An eye for an ear: zero-shot audio description leveraging an image captioner with audio-visual token distribution matchingHugo Malard, Michel Olvera, Stéphane Lathuilière, Slim EssidNeurIPS 2024 · 3 citations
- DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio SynthesisWenjie Tian, Xinfa Zhu, Haohe Liu, Zhixian Zhao et al.ACM MM 2025
- From Speaker to Dubber: Movie Dubbing with Prosody and Duration Consistency LearningZhedong Zhang, Liang Li, Gaoxiang Cong, Haibing Yin et al.ACM MM 2024 · 36 citations
- UniSpeech: Unified Speech Representation Learning with Labeled and Unlabeled DataChengyi Wang, Yu Wu, Yao Qian, Ken'ichi Kumatani et al.ICML 2021 · 140 citations
