MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis
Ho Kei Cheng, Masato Ishii, Akio Hayakawa, Takashi Shibuya, Alexander G. Schwing, Yuki Mitsufuji
摘要
We propose to synthesize high-quality and synchronized audio, given video and optional text conditions, using a novel multimodal joint training framework (MMAudio). In contrast to single-modality training conditioned on (limited) video data only, MMAudio is jointly trained with largerscale, readily available text-audio data to learn to generate semantically aligned high-quality audio samples. Additionally, we improve audio-visual synchrony with a conditional synchronization module that aligns video conditions with audio latents at the frame level. Trained with a flow matching objective, MMAudio achieves new video-to-audio stateof-the-art among public models in terms of audio quality, semantic alignment, and audio-visual synchronization, while having a low inference time (1.23s to generate an 8s clip) and just 157M parameters. MMAudio also achieves surprisingly competitive performance in text-to-audio generation, showing that joint training does not hinder single-modality performance. Code, models, and demo are available at: hkchengrex.github.io/MMAudio.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- VISTA: A Test-Time Self-Improving Video Generation AgentDo Xuan Long, Xingchen Wan, Hootan Nakhost, Chen-Yu Lee 等CVPR 2026 · 被引用 30 次
- Harmony: Harmonizing Audio and Video Generation through Cross-Task SynergyTeng Hu, Zhentao Yu, Guozhen Zhang, Zihan Su 等CVPR 2026 · 被引用 21 次
- T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video GenerationZhe Cao, Tao Wang, Jiaming Wang, Yanghai Wang 等ICML 2026 · 被引用 13 次
- Omni2Sound: Towards Unified Video-Text-to-Audio Generationyusheng dai, Zehua Chen, Yuxuan Jiang, Qiuhong Ke 等CVPR 2026 · 被引用 12 次
- EchoFoley: Event-Centric Hierarchical Control for Video Grounded Creative Sound GenerationBingxuan Li, Yiming Cui, Yicheng He, Yiwei Wang 等CVPR 2026 · 被引用 5 次
它引用的顶会 Paper20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
- AudioLDM: Text-to-Audio Generation with Latent Diffusion ModelsHaohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei 等ICML 2023 · 被引用 773 次
- Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion ModelsRongjie Huang, Jiawei Huang, Dongchao Yang, Yi Ren 等ICML 2023 · 被引用 469 次
- LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic AlignmentBin Zhu, Bin Lin, Munan Ning, Yang Yan 等ICLR 2024 · 被引用 403 次
相关 Paper
- TiVA: Time-Aligned Video-to-Audio GenerationXihua Wang, Yuyue Wang, Yihan Wu, Ruihua Song 等ACM MM 2024 · 被引用 6 次
- AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video GenerationMoayed Haji-Ali, Willi Menapace, Aliaksandr Siarohin, Ivan Skorokhodov 等ICCV 2025 · 被引用 3 次
- Read, Watch and Scream! Sound Generation from Text and VideoYujin Jeong, Yunji Kim, Sanghyuk Chun, Jiyoung LeeAAAI 2025 · 被引用 48 次
- Audio-Sync Video Generation with Multi-Stream Temporal ControlShuchen Weng, Haojie Zheng, Zheng Chang, Si Li 等NeurIPS 2025 · 被引用 14 次
- Aligning What Matters: Masked Latent Adaptation for Text-to-Audio-Video GenerationJiyang Zheng, Siqi Pan, Yu Yao, Zhaoqing Wang 等NeurIPS 2025 · 被引用 6 次
