Audio-Reasoner: Improving Reasoning Capability in Large Audio Language Models
Zhifei Xie, Mingbao Lin, Zihang Liu, Pengcheng Wu, Shuicheng Yan, Chunyan Miao
Abstract
Recent advancements in multimodal reasoning have largely overlooked the audio modality. We introduce Audio-Reasoner, a large-scale audio language model for deep reasoning in audio tasks. We meticulously curated a large-scale and diverse multi-task audio dataset with simple annotations. Then, we leverage closed-source models to conduct secondary labeling, QA generation, along with structured COT process. These datasets together form a high-quality reasoning dataset with 1.2 million reasoning-rich samples, which we name CoTA. Following inference scaling principles, we train Audio-Reasoner on CoTA, enabling it to achieve great logical capabilities in audio reasoning. Experiments show state-of-the-art performance across key benchmarks, including MMAU-mini (+25.42%), AIR-Bench chat/foundation(+14.57%/+10.13%), and MELD (+8.01%). Our findings stress the core of structured CoT training in advancing audio reasoning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 479c5975-a62e-4927-9a94-0a1a1aadd180Cited by top-tier papers19
- Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language ModelsSreyan Ghosh, Arushi Goel, Jaehyeon Kim, Sonal Kumar et al.NeurIPS 2025 · 299 citations
- STITCH: Simultaneous Thinking and Talking with Chunked Reasoning for Spoken Language ModelsCheng-Han Chiang, Xiaofei Wang, Linjie Li, Chung-Ching Lin et al.ICLR 2026 · 38 citations
- ParaS2S: Benchmarking and Aligning Spoken Language Models for Paralinguistic-aware Speech-to-Speech InteractionShu-Wen Yang, Ming Tu, Ting-Wei Liu, Xinghua Qu et al.ICLR 2026 · 29 citations
- MME-Emotion: A Holistic Evaluation Benchmark for Emotional Intelligence in Multimodal Large Language ModelsFan Zhang, Zebang Cheng, Chong Deng, Haoxuan Li et al.ICLR 2026 · 23 citations
- AudSemThinker: Enhancing Audio-Language Models Through Reasoning over Semantics of SoundGijs Wijngaard, Elia Formisano, Michele Esposito, Michel DumontierNeurIPS 2025 · 22 citations
Builds on20
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 1,792 citations
Related papers
- Mellow: a small audio language model for reasoningSoham Deshmukh, Satvik Dixit, Rita Singh, Bhiksha RajNeurIPS 2025 · 39 citations
- Incentivizing Consistent, Effective and Scalable Reasoning Capability in Audio LLMs via Reasoning Process RewardsJiajun Fan, Roger Ren, Jingyuan Li, Rahul Pandey et al.ICLR 2026 · 15 citations
- SoundMind: RL-Incentivized Logic Reasoning for Audio-Language ModelsXingjian Diao, Chunhui Zhang, Keyi Kong, Weiyi Wu et al.EMNLP 2025 · 1 citation
- Listen, Think, and UnderstandYuan Gong, Hongyin Luo, Alexander H. Liu, Leonid Karlinsky et al.ICLR 2024 · 247 citations
- Audio-Thinker: Guiding Large Audio Language Model When and How to Think via Reinforcement LearningShu Wu, Chenxing Li, Wenfu Wang, Hao Zhang et al.AAAI 2026 · 4 citations
