FlowDubber: Movie Dubbing with LLM-based Semantic-aware Learning and Flow Matching based Voice Enhancing
Gaoxiang Cong, Liang Li, Jiadong Pan, Zhedong Zhang, Amin Beheshti, Anton van den Hengel, Yuankai Qi, Qingming Huang
摘要
Movie Dubbing aims to convert scripts into speeches that align with the given movie clip in both temporal and emotional aspects while preserving the vocal timbre of a given brief reference audio. Existing methods focus primarily on reducing the word error rate while ignoring the importance of lip-sync and acoustic quality. To address these issues, we propose a novel dubbing architecture based on Large Language Model (LLM) and Conditional Flow Matching (CFM), named FlowDubber, which achieves high-quality audio-visual sync and pronunciation by incorporating a large speech language model with dual contrastive alignment while improving acoustic quality via Flow-based Voice Enhancing (FVE). First, we introduce Qwen2.5 as the backbone of large speech language model to learn the in-context sequence from movie scripts and reference audio. Second, the proposed semantic-aware learning focuses on capturing LLM semantic knowledge at the phoneme level, which facilitates mutual alignment with lip movement from silent video via Dual Contrastive Alignment (DCA). Third, the FVE introduces an LLM-based acoustics flow matching guidance to strengthen clarity by decoupling Classifier-Free Guidance (CFG) enhancement. Extensive experiments demonstrate that our method outperforms several state-of-the-art methods on two primary benchmarks. The demos are available at https://galaxycong.github.io/LLM-Flow-Dubber/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- UniAVGen: Unified Audio and Video Generation with Asymmetric Cross-Modal InteractionsGuozhen Zhang, Zixiang Zhou, Teng Hu, Ziqiao Peng 等CVPR 2026 · 被引用 40 次
- InstructDubber: Instruction-based Alignment for Zero-shot Movie DubbingZhedong Zhang, Liang Li, Gaoxiang Cong, Chunshan Liu 等AAAI 2026 · 被引用 3 次
- Hierarchical Codec Diffusion for Video-to-Speech GenerationJiaxin Ye, Gaoxiang Cong, Chenhui Wang, Xin-Cheng Wen 等CVPR 2026 · 被引用 3 次
- Expert-Teacher-Student Collaborative Learning for Domain Adaptive Object DetectionYiming Cui, Liang Li, Haibing Yin, Yuhan Gao 等CVPR 2026
它引用的顶会 Paper37
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 被引用 2,890 次
相关 Paper
- EmoDubber: Towards High Quality and Emotion Controllable Movie DubbingGaoxiang Cong, Jiadong Pan, Liang Li, Yuankai Qi 等CVPR 2025
- From Speaker to Dubber: Movie Dubbing with Prosody and Duration Consistency LearningZhedong Zhang, Liang Li, Gaoxiang Cong, Haibing Yin 等ACM MM 2024 · 被引用 36 次
- Towards Authentic Movie Dubbing with Retrieve-Augmented Director-Actor Interaction LearningRui Liu, Yuan Zhao, Zhenqi JiaAAAI 2026
- Prosody-Enhanced Acoustic Pre-training and Acoustic-Disentangled Prosody Adapting for Movie DubbingZhedong Zhang, Liang Li, Chenggang Yan, Chunshan Liu 等CVPR 2025
- Learning to Dub Movies via Hierarchical Prosody ModelsGaoxiang Cong, Liang Li, Yuankai Qi, Zheng-Jun Zha 等CVPR 2023
