Synchronized Video-to-Audio Generation via Mel Quantization-Continuum Decomposition
Juncheng Wang, Chao Xu, Cheng Yu, Lei Shang, Zhe Hu, Shujun Wang, Liefeng Bo
Abstract
Video-to-audio generation is essential for synthesizing realistic audio tracks that synchronize effectively with silent videos. Following the perspective of extracting essential signals from videos that can precisely control the mature text-to-audio generative diffusion models, this paper presents how to balance the representation of melspectrograms in terms of completeness and complexity through a new approach called Mel Quantization-Continuum Decomposition (Mel-QCD). We decompose the mel-spectrogram into three distinct types of signals, employing quantization or continuity to them, we can effectively predict them from video by a devised video-to-all (V2X) predictor. Then, the predicted signals are recomposed and fed into a ControlNet, along with a textual inversion design, to control the audio generation process. Our proposed Mel-QCD method demonstrates state-of-the-art performance across eight metrics, evaluating dimensions such as quality, synchronization, and semantic consistency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 42c656ca-f2be-420b-b6b9-eb2f46386eedCited by top-tier papers2
- SyncTrack: Rhythmic Stability and Synchronization in Multi-Track Music GenerationHongrui Wang, Fan Zhang, Zhiyuan Yu, Ziya Zhou et al.ICLR 2026
- Language Model Based Text-to-Audio Generation: Anti-Causally Aligned Collaborative Residual TransformersJuncheng Wang, Chao Xu, Cheng Yu, Zhe Hu et al.EMNLP 2025
Builds on19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- AudioLDM: Text-to-Audio Generation with Latent Diffusion ModelsHaohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei et al.ICML 2023 · 773 citations
Related papers
- Read, Watch and Scream! Sound Generation from Text and VideoYujin Jeong, Yunji Kim, Sanghyuk Chun, Jiyoung LeeAAAI 2025 · 48 citations
- OmniVDiff: Omni Controllable Video Diffusion for Generation and UnderstandingDianbing Xi, Jiepeng Wang, Yuanzhi Liang, Xi Qiu et al.AAAI 2026 · 14 citations
- Autoregressive Speech Synthesis without Vector QuantizationLingwei Meng, Long Zhou, Shujie Liu, Sanyuan Chen et al.ACL 2025 · 94 citations
- TiVA: Time-Aligned Video-to-Audio GenerationXihua Wang, Yuyue Wang, Yihan Wu, Ruihua Song et al.ACM MM 2024 · 6 citations
- Spatial-Temporal Decomposition and Alignment in Controllable Video-to-Music GenerationWeitao You, Heda Zuo, Junxian Wu, Dengming Zhang et al.ACM MM 2025
