Fine-grained Video Dubbing Duration Alignment with Segment Supervised Preference Optimization
Chaoqun Cui, Liangbin Huang, Shijing Wang, Zhe Tong, Zhaolong Huang, Xiao Zeng, Xiaofeng Liu
Abstract
Video dubbing aims to translate original speech in visual media programs from the source language to the target language, relying on neural machine translation and text-to-speech technologies. Due to varying information densities across languages, target speech often mismatches the source speech duration, causing audio-video synchronization issues that significantly impact viewer experience. In this study, we approach duration alignment in LLM-based video dubbing machine translation as a preference optimization problem. We propose the Segment Supervised Preference Optimization (SSPO) method, which employs a segmentwise sampling strategy and fine-grained loss to mitigate duration mismatches between source and target lines. Experimental results demonstrate that SSPO achieves superior performance in duration alignment tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Just-Dub-It: Video dubbing via Joint Audio-Visual DiffusionAnthony Chen, Naomi Ken Korem, Tavi Halperin, Matan Ben-Yosef et al.SIGGRAPH 2026 · 2 citations
- Hermes the Polyglot: A Unified Framework to Enhance Expressiveness for Multimodal Interlingual SubtitlingChaoqun Cui, Shijing Wang, Liangbin Huang, Qingqing Gu et al.WWW 2026 · 1 citation
- CineSRD: Leveraging Visual, Acoustic, and Linguistic Cues for Open-World Visual Media Speaker DiarizationLiangbin Huang, Xiaohua Liao, Chaoqun Cui, Shijing Wang et al.CVPR 2026
Builds on10
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- SimPO: Simple Preference Optimization with a Reference-Free RewardYu Meng, Mengzhou Xia, Danqi ChenNeurIPS 2024 · 1,203 citations
- Learning to summarize with human feedbackNisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel M. Ziegler et al.NeurIPS 2020 · 124 citations
Related papers
- VideoDubber: Machine Translation with Speech-Aware Length Control for Video DubbingYihan Wu, Junliang Guo, Xu Tan, Chen Zhang et al.AAAI 2023 · 33 citations
- InstructDubber: Instruction-based Alignment for Zero-shot Movie DubbingZhedong Zhang, Liang Li, Gaoxiang Cong, Chunshan Liu et al.AAAI 2026 · 3 citations
- From Speaker to Dubber: Movie Dubbing with Prosody and Duration Consistency LearningZhedong Zhang, Liang Li, Gaoxiang Cong, Haibing Yin et al.ACM MM 2024 · 36 citations
- From Utterance to Vividity: Training Expressive Subtitle Translation LLM via Adaptive Local Preference OptimizationChaoqun Cui, Shijing Wang, Liangbin Huang, Qingqing Gu et al.ICLR 2026 · 1 citation
- Neural Dubber: Dubbing for Videos According to ScriptsChenxu Hu, Qiao Tian, Tingle Li, Yuping Wang et al.NeurIPS 2021 · 62 citations
