Aurelius: Relation Aware Text-to-Audio Generation At Scale
Yuhang He, He Liang, Yash Jain, Andrew Markham, Vibhav Vineet
Abstract
We present Aurelius, a new framework that enables relation aware text-to-audio (TTA) generation research at scale. Given the lack of essential audio event and relation corpora, Aurelius contributes a large-scale audio event corpus AudioEventSet and another large-scale relation corpus AudioRelSet. Comprising 110 event categories, AudioEventSet maximally covers all commonly heard audio events and each event is unique, realistic and of high-quality. AudioRelSet consists of 100 relations, comprehensively covering the relations that present in the physical world or can be neatly described by text. As the two corpora provide audio event and relation independently, they can be combined to create massive <text,audio> pairs with our pair generation strategy to support relation aware TTA investigation at scale. We comprehensively benchmark all existing TTA models from both general and relation aware evaluation perspective. We further provide an in-depth investigation into scaling existing TTA models' relation aware generation by either training from scratch or leveraging cross-domain general TTA knowledge. The introduced corpora and the findings from investigation potentially facilitate future research on relation aware TTA generation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on14
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Score-based Generative Modeling in Latent SpaceArash Vahdat, Karsten Kreis, Jan KautzNeurIPS 2021 · 903 citations
- Simple and Controllable Music GenerationJade Copet, Felix Kreuk, Itai Gat, Tal Remez et al.NeurIPS 2023 · 843 citations
- AudioLDM: Text-to-Audio Generation with Latent Diffusion ModelsHaohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei et al.ICML 2023 · 773 citations
- Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion ModelsRongjie Huang, Jiawei Huang, Dongchao Yang, Yi Ren et al.ICML 2023 · 469 citations
Related papers
- RiTTA: Modeling Event Relations in Text-to-Audio GenerationYuhang He, Yash Jain, Xubo Liu, Andrew Markham et al.EMNLP 2025 · 1 citation
- TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference OptimizationChia-Yu Hung, Navonil Majumder, Zhifeng Kong, Ambuj Mehrish et al.ICLR 2026 · 67 citations
- AudioStory: Generating Long-Form Narrative Audio with Large Language ModelsYuxin Guo, Teng Wang, Yuying Ge, Shijie Ma et al.CVPR 2026 · 5 citations
- TTA-Bench: A Comprehensive Benchmark for Evaluating Text-to-Audio ModelsHui Wang, Cheng Liu, Junyang Chen, Haoze Liu et al.AAAI 2026 · 2 citations
- Audio-Reasoner: Improving Reasoning Capability in Large Audio Language ModelsZhifei Xie, Mingbao Lin, Zihang Liu, Pengcheng Wu et al.EMNLP 2025 · 5 citations
