V2C: Visual Voice Cloning
Qi Chen, Mingkui Tan, Yuankai Qi, Jiaqiu Zhou, Yuanqing Li, Qi Wu
Abstract
Existing Voice Cloning (VC) tasks aim to convert a para-graph text to a speech with desired voice specified by a ref-erence audio. This has significantly boosted the development of artificial speech applications. However, there also exist many scenarios that cannot be well reflected by these VC tasks, such as movie dubbing, which requires the speech to be with emotions consistent with the movie plots. To fill this gap, in this work we propose a new task named Vi-sual Voice Cloning (V2C), which seeks to convert a para-graph of text to a speech with both desired voice speci-fied by a reference audio and desired emotion specified by a reference video. To facilitate research in this field, we construct a dataset, V2C-Animation, and propose a strong baseline based on existing state-of-the-art (SoTA) VC techniques. Our dataset contains 10,217 animated movie clips covering a large variety of genres (e.g., Comedy, Fantasy) and emotions (e.g., happy, sad). We further design a set of evaluation metrics, named MCD-DTW-SL, which help eval-uate the similarity between ground-truth speeches and the synthesised ones. Extensive experimental results show that even SoTA VC methods cannot generate satisfying speeches for our V2C task. We hope the proposed new task together with the constructed dataset and evaluation metric will fa-cilitate the research in the field of voice cloning and broader vision-and-language community. Source code and dataset will be released in https://github.com/chenqi008/V2C.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 20feda37-67a6-4bbd-b905-43e53bddab54Cited by top-tier papers8
- InstructDubber: Instruction-based Alignment for Zero-shot Movie DubbingZhedong Zhang, Liang Li, Gaoxiang Cong, Chunshan Liu et al.AAAI 2026 · 3 citations
- Hierarchical Codec Diffusion for Video-to-Speech GenerationJiaxin Ye, Gaoxiang Cong, Chenhui Wang, Xin-Cheng Wen et al.CVPR 2026 · 3 citations
- Learning to Dub Movies via Hierarchical Prosody ModelsGaoxiang Cong, Liang Li, Yuankai Qi, Zheng-Jun Zha et al.CVPR 2023
- Prosody-Enhanced Acoustic Pre-training and Acoustic-Disentangled Prosody Adapting for Movie DubbingZhedong Zhang, Liang Li, Chenggang Yan, Chunshan Liu et al.CVPR 2025
- Emotional Face-to-SpeechJiaxin Ye, Boyuan Cao, Hongming ShanICML 2025
Builds on5
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 2,890 citations
- VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and TextHassan Akbari, Liangzhe Yuan, Rui Qian, Wei-Hong Chuang et al.NeurIPS 2021 · 782 citations
- FastSpeech 2: Fast and High-Quality End-to-End Text to SpeechYi Ren, Chenxu Hu, Xu Tan, Tao Qin et al.ICLR 2021 · 513 citations
- COOT: Cooperative Hierarchical Transformer for Video-Text Representation LearningSimon Ging, Mohammadreza Zolfaghari, Hamed Pirsiavash, Thomas BroxNeurIPS 2020 · 186 citations
- AdaSpeech: Adaptive Text to Speech for Custom VoiceMingjian Chen, Xu Tan, Bohan Li, Yanqing Liu et al.ICLR 2021 · 79 citations
Related papers
- EmoDubber: Towards High Quality and Emotion Controllable Movie DubbingGaoxiang Cong, Jiadong Pan, Liang Li, Yuankai Qi et al.CVPR 2025
- From Speaker to Dubber: Movie Dubbing with Prosody and Duration Consistency LearningZhedong Zhang, Liang Li, Gaoxiang Cong, Haibing Yin et al.ACM MM 2024 · 36 citations
- VoiceCraft-Dub: Automated Video Dubbing with Neural Codec Language ModelsSung-Bin Kim, Jeongsoo Choi, Puyuan Peng, Joon Son Chung et al.ICCV 2025
- Neural Dubber: Dubbing for Videos According to ScriptsChenxu Hu, Qiao Tian, Tingle Li, Yuping Wang et al.NeurIPS 2021 · 62 citations
- Towards Authentic Movie Dubbing with Retrieve-Augmented Director-Actor Interaction LearningRui Liu, Yuan Zhao, Zhenqi JiaAAAI 2026
