ADIFF: Explaining audio difference using natural language
Soham Deshmukh, Shuo Han, Rita Singh, Bhiksha Raj
Abstract
Understanding and explaining differences between audio recordings is crucial for fields like audio forensics, quality assessment, and audio generation. This involves identifying and describing audio events, acoustic scenes, signal characteristics, and their emotional impact on listeners. This paper stands out as the first work to comprehensively study the task of explaining audio differences and then propose benchmark, baselines for the task. First, we present two new datasets for audio difference explanation derived from the AudioCaps and Clotho audio captioning datasets. Using Large Language Models (LLMs), we generate three levels of difference explanations: (1) concise descriptions of audio events and objects, (2) brief sentences about audio events, acoustic scenes, and signal properties, and (3) comprehensive explanations that include semantics and listener emotions. For the baseline, we use prefix tuning where audio embeddings from two audio files are used to prompt a frozen language model. Our empirical analysis and ablation studies reveal that the naive baseline struggles to distinguish perceptually similar sounds and generate detailed tier 3 explanations. To address these limitations, we propose ADIFF, which introduces a cross-projection module, position captioning, and a three-step training process to enhance the model's ability to produce detailed explanations. We evaluate our model using objective metrics and human evaluation and show our model enhancements lead to significant improvements in performance over naive baseline and SoTA Audio-Language Model (ALM) Qwen Audio. Lastly, we conduct multiple ablation studies to study the effects of cross-projection, language model parameters, position captioning, third stage fine-tuning, and present our findings. Our benchmarks, findings, and strong baseline pave the way for nuanced and human-like explanations of audio differences.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Mellow: a small audio language model for reasoningSoham Deshmukh, Satvik Dixit, Rita Singh, Bhiksha RajNeurIPS 2025 · 39 citations
- LLM2Fx-Tools: Tool Calling for Music Post-ProductionSeungHeon Doh, Junghyun Koo, Marco A. Martínez-Ramírez, Woosung Choi et al.ICLR 2026 · 10 citations
- Towards Holistic Evaluation of Large Audio-Language Models: A Comprehensive SurveyChih-Kai Yang, Neo S. Ho, Hung-yi LeeEMNLP 2025 · 7 citations
Builds on15
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- AudioLDM: Text-to-Audio Generation with Latent Diffusion ModelsHaohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei et al.ICML 2023 · 773 citations
- SALMONN: Towards Generic Hearing Abilities for Large Language ModelsChangli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen et al.ICLR 2024 · 557 citations
- Pengi: An Audio Language Model for Audio TasksSoham Deshmukh, Benjamin Elizalde, Rita Singh, Huaming WangNeurIPS 2023 · 352 citations
Related papers
- Learning to See through Sound: From VggCaps to Multi2Cap for Richer Automated Audio CaptioningSangyeon Cho, Mingi Kim, Jinkwon Hwang, Jaehoon Go et al.EMNLP 2025
- Audio Entailment: Assessing Deductive Reasoning for Audio UnderstandingSoham Deshmukh, Shuo Han, Hazim T. Bukhari, Benjamin Elizalde et al.AAAI 2025 · 23 citations
- Synthio: Augmenting Small-Scale Audio Classification Datasets with Synthetic DataSreyan Ghosh, Sonal Kumar, Zhifeng Kong, Rafael Valle et al.ICLR 2025
- ETTA: Elucidating the Design Space of Text-to-Audio ModelsSang-gil Lee, Zhifeng Kong, Arushi Goel, Sungwon Kim et al.ICML 2025
- DIFFA: Large Language Diffusion Models Can Listen and UnderstandJiaming Zhou, Hongjie Chen, Shiwan Zhao, Jian Kang et al.AAAI 2026 · 10 citations
