Prosody-Enhanced Acoustic Pre-training and Acoustic-Disentangled Prosody Adapting for Movie Dubbing
Zhedong Zhang, Liang Li, Chenggang Yan, Chunshan Liu, Anton van den Hengel, Yuankai Qi
摘要
Movie dubbing describes the process of transforming a script into speech that aligns temporally and emotionally with a given movie clip while exemplifying the speaker's voice demonstrated in a short reference audio clip. This task demands the model bridge character performances and complicated prosody structures to build a high-quality video-synchronized dubbing track. The limited scale of movie dubbing datasets, along with the background noise inherent in audio data, hinder the acoustic modeling performance of trained models. To address these issues, we propose an acoustic-prosody disentangled two-stage method to achieve high-quality dubbing generation with precise prosody alignment. First, we propose a prosody-enhanced acoustic pre-training to develop robust acoustic modeling capabilities. Then, we freeze the pre-trained acoustic system and design an acoustic-disentangled framework to model prosodic text features and dubbing style while maintaining acoustic quality. Additionally, we incorporate an in-domain emotion analysis module to reduce the impact of visual domain shifts across different movies, thereby enhancing emotion-prosody alignment. Extensive experiments show that our method performs favorably against the stateof-the-art models on two primary benchmarks. The demos and source code are available at here.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- InstructDubber: Instruction-based Alignment for Zero-shot Movie DubbingZhedong Zhang, Liang Li, Gaoxiang Cong, Chunshan Liu 等AAAI 2026 · 被引用 3 次
- Progressive Homeostatic and Plastic Prompt Tuning for Audio-Visual Multi-Task Incremental LearningJiong Yin, Liang Li, Jiehua Zhang, Yuhan Gao 等ICCV 2025 · 被引用 3 次
- FlowDubber: Movie Dubbing with LLM-based Semantic-aware Learning and Flow Matching based Voice EnhancingGaoxiang Cong, Liang Li, Jiadong Pan, Zhedong Zhang 等ACM MM 2025 · 被引用 2 次
- Debiased Teacher for Day-to-Night Domain Adaptive Object DetectionYiming Cui, Liang Li, Haibing Yin, Yuhan Gao 等ICCV 2025 · 被引用 2 次
- Temporal Calibrating and Distilling for Scene-Text Aware Text-Video RetrievalZhiqian Zhao, Liang Li, Lei Shen, Xichun Sheng 等AAAI 2026 · 被引用 1 次
它引用的顶会 Paper23
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 被引用 2,890 次
- Grad-TTS: A Diffusion Probabilistic Model for Text-to-SpeechVadim Popov, Ivan Vovk, Vladimir Gogoryan, Tasnima Sadekova 等ICML 2021 · 被引用 715 次
- Glow-TTS: A Generative Flow for Text-to-Speech via Monotonic Alignment SearchJaehyeon Kim, Sungwon Kim, Jungil Kong, Sungroh YoonNeurIPS 2020 · 被引用 663 次
- YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot Voice Conversion for EveryoneEdresson Casanova, Julian Weber, Christopher Dane Shulby, Arnaldo Cândido Júnior 等ICML 2022 · 被引用 602 次
相关 Paper
- From Speaker to Dubber: Movie Dubbing with Prosody and Duration Consistency LearningZhedong Zhang, Liang Li, Gaoxiang Cong, Haibing Yin 等ACM MM 2024 · 被引用 36 次
- Learning to Dub Movies via Hierarchical Prosody ModelsGaoxiang Cong, Liang Li, Yuankai Qi, Zheng-Jun Zha 等CVPR 2023
- EmoDubber: Towards High Quality and Emotion Controllable Movie DubbingGaoxiang Cong, Jiadong Pan, Liang Li, Yuankai Qi 等CVPR 2025
- V2C: Visual Voice CloningQi Chen, Mingkui Tan, Yuankai Qi, Jiaqiu Zhou 等CVPR 2022
- Towards Authentic Movie Dubbing with Retrieve-Augmented Director-Actor Interaction LearningRui Liu, Yuan Zhao, Zhenqi JiaAAAI 2026
