Tri-Ergon: Fine-Grained Video-to-Audio Generation with Multi-Modal Conditions and LUFS Control
Bingliang Li, Fengyu Yang, Yuxin Mao, Qingwen Ye, Hongkai Chen, Yiran Zhong
摘要
Video-to-audio (V2A) generation utilizes visual-only video features to produce realistic sounds that correspond to the scene. However, current V2A models often lack fine-grained control over the generated audio, especially in terms of loudness variation and the incorporation of multi-modal conditions. To overcome these limitations, we introduce Tri-Ergon, a diffusion-based V2A model that incorporates textual, auditory, and pixel-level visual prompts to enable detailed and semantically rich audio synthesis. Additionally, we introduce Loudness Units relative to Full Scale (LUFS) embedding, which allows for precise manual control of the loudness changes over time for individual audio channels, enabling our model to effectively address the intricate correlation of video and audio in real-world Foley workflows. Tri-Ergon is capable of creating 44.1 kHz high-fidelity stereo audio clips of varying lengths up to 60 seconds, which significantly outperforms existing state-of-the-art V2A methods that typically generate mono audio for a fixed duration.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- OmniSonic: Towards Universal and Holistic Audio Generation from Video and TextWeiguo Pian, Saksham Singh Kushwaha, Zhimin Chen, Shijian Deng 等CVPR 2026 · 被引用 2 次
- Towards Open-Vocabulary Audio-Visual Event LocalizationJinxing Zhou, Dan Guo, Ruohao Guo, Yuxin Mao 等CVPR 2025
它引用的顶会 Paper13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic AlignmentBin Zhu, Bin Lin, Munan Ning, Yang Yan 等ICLR 2024 · 被引用 403 次
- VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and DatasetSihan Chen, Handong Li, Qunbo Wang, Zijia Zhao 等NeurIPS 2023 · 被引用 246 次
- Fast Timing-Conditioned Latent Audio DiffusionZach Evans, CJ Carr, Josiah Taylor, Scott H. Hawley 等ICML 2024 · 被引用 220 次
- Diff-Foley: Synchronized Video-to-Audio Synthesis with Latent Diffusion ModelsSimian Luo, Chuanhao Yan, Chenxu Hu, Hang ZhaoNeurIPS 2023 · 被引用 192 次
相关 Paper
- SALSA-V: Shortcut-Augmented Long-form Synchronized Audio from VideosAmir Dellali, Luca Lanzendörfer, Florian Grötschla, Roger WattenhoferICML 2026
- CAFA: A Controllable Automatic Foley ArtistRoi Benita, Michael Finkelson, Tavi Halperin, Gleb Sterkin 等ICCV 2025 · 被引用 1 次
- Read, Watch and Scream! Sound Generation from Text and VideoYujin Jeong, Yunji Kim, Sanghyuk Chun, Jiyoung LeeAAAI 2025 · 被引用 48 次
- AC-Foley: Reference-Audio-Guided Video-to-Audio Synthesis with Acoustic TransferPengjun Fang, Yingqing He, Yazhou Xing, Qifeng Chen 等ICLR 2026 · 被引用 3 次
- Hear What You See: Video-to-Audio Generation with Diffusion Transformer and Semantic-Temporal Alignment-Ranked Direct Preference OptimizationKai Wang, Tao Zhou, Jiayi Lei, Jing Wang 等CVPR 2026
