DiVAS: Video and Audio Synchronization with Dynamic Frame Rates
Clara Fernandez-Labrador, Mertcan Akçay, Eitan Abecassis, Joan Massich, Christopher Schroers
Abstract
Synchronization issues between audio and video are one of the most disturbing quality defects in film production and live broadcasting. Even a discrepancy as short as 45 milliseconds can degrade the viewer's experience enough to warrant manual quality checks over entire movies. In this paper, we study the automatic discovery of such issues. Specifically, we focus on the alignment of lip movements with spoken words, targeting realistic production scenarios which can include background noise and music, intricate head poses, excessive makeup, or scenes with multiple individuals where the speaker is unknown. Our model's robustness also extends to various media specifications, including different video frame rates and audio sample rates. To address these challenges, we present a model fully based on Transformers that encodes face crops or full video frames and raw audio using timestamp information, identifies the speaker and provides highly accurate synchronization predictions much faster than previous methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8f7a9626-aafb-4d13-9c50-91638fc016bcCited by top-tier papers1
Ask how each one uses itBuilds on2
- VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and TextHassan Akbari, Liangzhe Yuan, Rui Qian, Wei-Hong Chuang et al.NeurIPS 2021 · 782 citations
- Is Someone Speaking?: Exploring Long-term Temporal Features for Audio-visual Active Speaker DetectionRuijie Tao, Zexu Pan, Rohan Kumar Das, Xinyuan Qian et al.ACM MM 2021 · 154 citations
Related papers
- OmniSync: Towards Universal Lip Synchronization via Diffusion TransformersZiqiao Peng, Jiwen Liu, Haoxian Zhang, Xiaoqiang Liu et al.NeurIPS 2025 · 30 citations
- Identity-Preserving Talking Face Generation with Landmark and Appearance PriorsWeizhi Zhong, Chaowei Fang, Yinqi Cai, Pengxu Wei et al.CVPR 2023
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 869 citations
- VisualVoice: Audio-Visual Speech Separation With Cross-Modal ConsistencyRuohan Gao, Kristen GraumanCVPR 2021
- FantasyTalking: Realistic Talking Portrait Generation via Coherent Motion SynthesisMengchao Wang, Qiang Wang, Fan Jiang, Yaqi Fan et al.ACM MM 2025 · 9 citations
