AudioGenX: Explainability on Text-to-Audio Generative Models
Hyunju Kang, Geonhee Han, Yoonjae Jeong, Hogun Park
摘要
Text-to-audio generation models (TAG) have achieved significant advances in generating audio conditioned on text descriptions. However, a critical challenge lies in the lack of transparency regarding how each textual input impacts the generated audio. To address this issue, we introduce Audio-GenX, an Explainable AI (XAI) method that provides explanations for text-to-audio generation models by highlighting the importance of input tokens. AudioGenX optimizes an Explainer by leveraging factual and counterfactual objective functions to provide faithful explanations at the audio token level. This method offers a detailed and comprehensive understanding of the relationship between text inputs and audio outputs, enhancing both the explainability and trustworthiness of TAG models. Extensive experiments demonstrate the effectiveness of AudioGenX in producing faithful explanations, benchmarked against existing methods using novel evaluation metrics specifically designed for audio generation tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- AudioLDM: Text-to-Audio Generation with Latent Diffusion ModelsHaohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei 等ICML 2023 · 被引用 773 次
- On Explainability of Graph Neural Networks via Subgraph ExplorationsHao Yuan, Haiyang Yu, Jie Wang, Kang Li 等ICML 2021 · 被引用 498 次
- Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion ModelsRongjie Huang, Jiawei Huang, Dongchao Yang, Yi Ren 等ICML 2023 · 被引用 469 次
- Generic Attention-model Explainability for Interpreting Bi-Modal and Encoder-Decoder TransformersHila Chefer, Shir Gur, Lior WolfICCV 2021 · 被引用 451 次
- Prompt-to-Prompt Image Editing with Cross-Attention ControlAmir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman 等ICLR 2023 · 被引用 361 次
相关 Paper
- AudioGen: Textually Guided Audio GenerationFelix Kreuk, Gabriel Synnaeve, Adam Polyak, Uriel Singer 等ICLR 2023 · 被引用 54 次
- Listenable Maps for Zero-Shot Audio ClassifiersFrancesco Paissan, Luca Della Libera, Mirco Ravanelli, Cem SubakanNeurIPS 2024 · 被引用 5 次
- TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference OptimizationChia-Yu Hung, Navonil Majumder, Zhifeng Kong, Ambuj Mehrish 等ICLR 2026 · 被引用 67 次
- AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video GenerationZiwei Zhou, Zeyuan Lai, Rui Wang, Yifan Yang 等ICML 2026 · 被引用 8 次
- TTA-Bench: A Comprehensive Benchmark for Evaluating Text-to-Audio ModelsHui Wang, Cheng Liu, Junyang Chen, Haoze Liu 等AAAI 2026 · 被引用 2 次
