MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark
S. Sakshi, Utkarsh Tyagi, Sonal Kumar, Ashish Seth, Ramaneswaran Selvakumar, Oriol Nieto, Ramani Duraiswami, Sreyan Ghosh, Dinesh Manocha
Abstract
The ability to comprehend audio—which includes speech, non-speech sounds, and music—is crucial for AI agents to interact effectively with the world. We present MMAU, a novel benchmark designed to evaluate multimodal audio understanding models on tasks requiring expert-level knowledge and complex reasoning. MMAU comprises 10k carefully curated audio clips paired with human-annotated natural language questions and answers spanning speech, environmental sounds, and music. It includes information extraction and reasoning questions, requiring models to demonstrate 27 distinct skills across unique and challenging tasks. Unlike existing benchmarks, MMAU emphasizes advanced perception and reasoning with domain-specific knowledge, challenging models to tackle tasks akin to those faced by experts. We assess 18 open-source and proprietary (Large) Audio-Language Models, demonstrating the significant challengesposed by MMAU. Notably, even the most advanced Gemini 2.0 Flash achieves only 59.93% accuracy, and the state-of-the-art open-source Qwen2-Audio achieves only 52.50%, highlighting considerable room for improvement. We believe MMAU will drive the audio and multimodal research community to develop more advanced audio understanding models capable of solving complex audio tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fe3de1fd-68bf-4613-a149-0881f4ff2462Cited by top-tier papers46
- Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language ModelsSreyan Ghosh, Arushi Goel, Jaehyeon Kim, Sonal Kumar et al.NeurIPS 2025 · 299 citations
- MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning BenchmarkDingdong Wang, Junan Li, Jincenzi Wu, Dongchao Yang et al.ICLR 2026 · 143 citations
- Mellow: a small audio language model for reasoningSoham Deshmukh, Satvik Dixit, Rita Singh, Bhiksha RajNeurIPS 2025 · 39 citations
- Omni-Captioner: Data Pipeline, Models, and Benchmark for Omni Detailed PerceptionZiyang Ma, Ruiyang Xu, Zhenghao Xing, Yunfei Chu et al.ICLR 2026 · 36 citations
- Music Flamingo: Scaling Music Understanding in Audio Language ModelsSreyan Ghosh, Arushi Goel, Lasha Koroshinadze, Sang-gil Lee et al.ICLR 2026 · 33 citations
Builds on13
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- Listen, Think, and UnderstandYuan Gong, Hongyin Luo, Alexander H. Liu, Leonid Karlinsky et al.ICLR 2024 · 247 citations
- Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue AbilitiesZhifeng Kong, Arushi Goel, Rohan Badlani, Wei Ping et al.ICML 2024 · 207 citations
Related papers
- MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General IntelligenceSonal Kumar, Simon Sedlácek, Vaibhavi Lokegaonkar, Fernando López et al.AAAI 2026 · 1 citation
- MAVERIX: Multimodal Audio-Visual Evaluation and Recognition IndeXLiuyue Xie, Avik Kuthiala, George Z. Wei, Ce Zheng et al.AAAI 2026 · 1 citation
- SciVideoBench: Benchmarking Scientific Video Reasoning in Large Multimodal ModelsAndong Deng, Taojiannan Yang, Shoubin Yu, Lincoln Spencer et al.ICML 2026 · 7 citations
- MMMU: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert AGIXiang Yue, Yuansheng Ni, Tianyu Zheng, Kai Zhang et al.CVPR 2024 · 213 citations
- MMVU: Measuring Expert-Level Multi-Discipline Video UnderstandingYilun Zhao, Haowei Zhang, Lujing Xie, Tongyan Hu et al.CVPR 2025
