Audio Generation with Multiple Conditional Diffusion Model
Zhifang Guo, Jianguo Mao, Rui Tao, Long Yan, Kazushige Ouchi, Hong Liu, Xiangdong Wang
摘要
Text-based audio generation models have limitations as they cannot encompass all the information in audio, leading to restricted controllability when relying solely on text. To address this issue, we propose a novel model that enhances the controllability of existing pre-trained text-to-audio models by incorporating additional conditions including content (timestamp) and style (pitch contour and energy contour) as supplements to the text. This approach achieves fine-grained control over the temporal order, pitch, and energy of generated audio. To preserve the diversity of generation, we employ a trainable control condition encoder that is enhanced by a large language model and a trainable Fusion-Net to encode and fuse the additional conditions while keeping the weights of the pre-trained text-to-audio model frozen. Due to the lack of suitable datasets and evaluation metrics, we consolidate existing datasets into a new dataset comprising the audio and corresponding conditions and use a series of evaluation metrics to evaluate the controllability performance. Experimental results demonstrate that our model successfully achieves finegrained control to accomplish controllable audio generation. Audio samples and our dataset are publicly available 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Training-free Mixed-Resolution Latent Upsampling for Spatially Accelerated Diffusion TransformersWongi Jeong, Kyungryeol Lee, Hoigi Seo, Se Young ChunCVPR 2026 · 被引用 10 次
- ControlAudio: Tackling Text-Guided, Timing-Indicated and Intelligible Audio Generation via Progressive Diffusion ModelingYuxuan Jiang, Zehua Chen, Zeqian Ju, Yusheng Dai 等ACL 2026 · 被引用 8 次
- FIND: Fine-tuning Initial Noise Distribution with Policy Optimization for Diffusion ModelsChanggu Chen, Libing Yang, Xiaoyan Yang, Lianggangxu Chen 等ACM MM 2024 · 被引用 8 次
- FreeAudio: Training-Free Timing Planning for Controllable Long-Form Text-to-Audio GenerationYuxuan Jiang, Zehua Chen, Zeqian Ju, Chang Li 等ACM MM 2025 · 被引用 5 次
- Hear What Matters! Text-conditioned Selective Video-to-Audio GenerationJunwon Lee, Juhan Nam, Jiyoung LeeCVPR 2026 · 被引用 4 次
它引用的顶会 Paper9
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam 等ICML 2022 · 被引用 4,691 次
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 被引用 2,890 次
- AudioLDM: Text-to-Audio Generation with Latent Diffusion ModelsHaohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei 等ICML 2023 · 被引用 773 次
- Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion ModelsRongjie Huang, Jiawei Huang, Dongchao Yang, Yi Ren 等ICML 2023 · 被引用 469 次
相关 Paper
- Read, Watch and Scream! Sound Generation from Text and VideoYujin Jeong, Yunji Kim, Sanghyuk Chun, Jiyoung LeeAAAI 2025 · 被引用 48 次
- ControlSpeech: Towards Simultaneous and Independent Zero-shot Speaker Cloning and Zero-shot Language Style ControlShengpeng Ji, Qian Chen, Wen Wang, Jialong Zuo 等ACL 2025
- Temporal-Conditioned Symbolic Alignment for Controllable Text-to-Music GenerationZihao Zhang, Xingjiao Wu, Junjie Xu, Tianlong Ma 等ACM MM 2025
- MuseControlLite: Multifunctional Music Generation with Lightweight ConditionersFang-Duo Tsai, Shih-Lun Wu, Weijaw Lee, Sheng-Ping Yang 等ICML 2025
- AudioX: A Unified Framework for Anything-to-Audio GenerationZeyue Tian, Zhaoyang Liu, Yizhu Jin, Ruibin Yuan 等ICLR 2026 · 被引用 38 次
