Read, Watch and Scream! Sound Generation from Text and Video
Yujin Jeong, Yunji Kim, Sanghyuk Chun, Jiyoung Lee
Abstract
Despite the impressive progress of multimodal generative models, video-to-audio generation still suffers from limited performance and limits the flexibility to prioritize sound synthesis for specific objects within the scene. Conversely, text-to-audio generation methods generate high-quality audio but pose challenges in ensuring comprehensive scene depiction and time-varying control. To tackle these challenges, we propose a novel video-and-text-to-audio generation method, called ReWaS, where video serves as a conditional control for a text-to-audio generation model. Especially, our method estimates the structural information of sound (namely, energy) from the video while receiving key content cues from a user prompt. We employ a well-performing text-to-audio model to consolidate the video control, which is much more efficient for training multimodal diffusion models with massive triplet-paired (audio-video-text) data. In addition, by separating the generative components of audio, it becomes a more flexible system that allows users to freely adjust the energy, surrounding environment, and primary sound source according to their preferences. Experimental results demonstrate that our method shows superiority in terms of quality, controllability, and training efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4a2661ea-88d7-4ed4-b673-c3268cfd1be4Cited by top-tier papers10
- UniAVGen: Unified Audio and Video Generation with Asymmetric Cross-Modal InteractionsGuozhen Zhang, Zixiang Zhou, Teng Hu, Ziqiao Peng et al.CVPR 2026 · 40 citations
- JavisDiT++: Unified Modeling and Optimization for Joint Audio-Video GenerationKai Liu, Yanhao Zheng, Kai Wang, Shengqiong Wu et al.ICLR 2026 · 24 citations
- Hear What Matters! Text-conditioned Selective Video-to-Audio GenerationJunwon Lee, Juhan Nam, Jiyoung LeeCVPR 2026 · 4 citations
- AC-Foley: Reference-Audio-Guided Video-to-Audio Synthesis with Acoustic TransferPengjun Fang, Yingqing He, Yazhou Xing, Qifeng Chen et al.ICLR 2026 · 3 citations
- PAVAS: Physics-Aware Video-to-Audio SynthesisOh Hyun-Bin, Yuhta Takida, Toshimitsu Uesaka, Tae-Hyun Oh et al.CVPR 2026 · 2 citations
Builds on21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 2,890 citations
Related papers
- MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio SynthesisHo Kei Cheng, Masato Ishii, Akio Hayakawa, Takashi Shibuya et al.CVPR 2025
- Tell What You Hear From What You See - Video to Audio Generation Through TextXiulong Liu, Kun Su, Eli ShlizermanNeurIPS 2024 · 46 citations
- Aligning What Matters: Masked Latent Adaptation for Text-to-Audio-Video GenerationJiyang Zheng, Siqi Pan, Yu Yao, Zhaoqing Wang et al.NeurIPS 2025 · 6 citations
- TiVA: Time-Aligned Video-to-Audio GenerationXihua Wang, Yuyue Wang, Yihan Wu, Ruihua Song et al.ACM MM 2024 · 6 citations
- OmniVDiff: Omni Controllable Video Diffusion for Generation and UnderstandingDianbing Xi, Jiepeng Wang, Yuanzhi Liang, Xi Qiu et al.AAAI 2026 · 14 citations
