Seeing and Hearing: Open-domain Visual-Audio Generation with Diffusion Latent Aligners
Yazhou Xing, Yingqing He, Zeyue Tian, Xintao Wang, Qifeng Chen
摘要
Video and audio content creation serves as the core technique for the movie industry and professional users. Re-cently, existing diffusion-based methods tackle video and audio generation separately, which hinders the technique transfer from academia to industry. In this work, we aim at filling the gap, with a carefully designed optimization-based framework for cross-visual-audio and joint-visual-audio generation. We observe the powerful generation abil-ity of off-the-shelf video or audio generation models. Thus, instead of training the giant models from scratch, we pro-pose to bridge the existing strong models with a shared la-tent representation space. Specifically, we propose a mul-timodality latent aligner with the pre-trained ImageBind model. Our latent aligner shares a similar core as the clas-sifier guidance that guides the diffusion denoising process during inference time. Through carefully designed opti-mization strategy and loss functions, we show the superior performance of our method on joint video-audio generation, visual-steered audio generation, and audio-steered vi-sual generation tasks. The project website can be found at https://yzxing87.github.io/Seeing-and-Hearing/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper57
- JavisDiT: Joint Audio-Video Diffusion Transformer with Hierarchical Spatio-Temporal Prior SynchronizationKai Liu, Wei Li, Lai Chen, Shengqiong Wu 等ICLR 2026 · 被引用 89 次
- Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow MatchingYongqi Wang, Wenxiang Guo, Rongjie Huang, Jiawei Huang 等NeurIPS 2024 · 被引用 73 次
- Read, Watch and Scream! Sound Generation from Text and VideoYujin Jeong, Yunji Kim, Sanghyuk Chun, Jiyoung LeeAAAI 2025 · 被引用 48 次
- UniAVGen: Unified Audio and Video Generation with Asymmetric Cross-Modal InteractionsGuozhen Zhang, Zixiang Zhou, Teng Hu, Ziqiao Peng 等CVPR 2026 · 被引用 40 次
- AudioX: A Unified Framework for Anything-to-Audio GenerationZeyue Tian, Zhaoyang Liu, Yizhu Jin, Ruibin Yuan 等ICLR 2026 · 被引用 38 次
它引用的顶会 Paper20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam 等ICML 2022 · 被引用 4,691 次
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan 等NeurIPS 2022 · 被引用 2,948 次
相关 Paper
- AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video GenerationMoayed Haji-Ali, Willi Menapace, Aliaksandr Siarohin, Ivan Skorokhodov 等ICCV 2025 · 被引用 3 次
- Aligning What Matters: Masked Latent Adaptation for Text-to-Audio-Video GenerationJiyang Zheng, Siqi Pan, Yu Yao, Zhaoqing Wang 等NeurIPS 2025 · 被引用 6 次
- MMDisCo: Multi-Modal Discriminator-Guided Cooperative Diffusion for Joint Audio and Video GenerationAkio Hayakawa, Masato Ishii, Takashi Shibuya, Yuki MitsufujiICLR 2025
- Animate and Sound an ImageXihua Wang, Ruihua Song, Chongxuan Li, Xin Cheng 等CVPR 2025
- Customized Condition Controllable Generation for Video SoundtrackFan Qi, Kunsheng Ma, Changsheng XuCVPR 2025
