MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark
Dingdong Wang, Junan Li, Jincenzi Wu, Dongchao Yang, Xueyuan Chen, Tianhua Zhang, Helen Meng
Abstract
Speech inherently contains rich acoustic information that extends far beyond the textual language. In real-world spoken communication, effective interpretation often requires integrating semantic meaning (e.g., content), paralinguistic features (e.g., emotions, speed, pitch) and phonological characteristics (e.g., prosody, intonation, rhythm), which are embedded in speech. While recent multimodal Speech Large Language Models (SpeechLLMs) have demonstrated remarkable capabilities in processing audio, their ability to perform fine-grained perception and complex reasoning in natural speech remains largely unexplored. To address this gap, we introduce MMSU, a comprehensive benchmark designed specifically for understanding and reasoning in speech. MMSU comprises 5,000 meticulously curated audio-question-answer triplets across 47 distinct tasks. Notably, linguistic theory forms the foundation of speech language understanding (SLU), yet existing benchmarks have paid insufficient attention to this fundamental aspect and fail to capture the broader linguistic picture. To ground our benchmark in linguistic principles, we systematically incorporate a wide range of linguistic phenomena, including phonetics, prosody, rhetoric, syntactics, semantics, and paralinguistics. Through a rigorous evaluation of 22 advanced SpeechLLMs, we identify substantial room for improvement in existing models. MMSU establishes a new standard for comprehensive assessment of SLLU, providing valuable insights for developing more sophisticated human-AI speech interaction systems. MMSU benchmark is available at https://huggingface.co/datasets/ddwang2000/MMSU.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 09332584-ad16-428c-a653-6df40f648d59Cited by top-tier papers11
- Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language ModelsSreyan Ghosh, Arushi Goel, Jaehyeon Kim, Sonal Kumar et al.NeurIPS 2025 · 299 citations
- Incentivizing Consistent, Effective and Scalable Reasoning Capability in Audio LLMs via Reasoning Process RewardsJiajun Fan, Roger Ren, Jingyuan Li, Rahul Pandey et al.ICLR 2026 · 15 citations
- EchoMind: An Interrelated Multi-level Benchmark for Evaluating Empathetic Speech Language ModelsLi Zhou, Lutong Yu, You Lyu, Yihang Lin et al.ICLR 2026 · 13 citations
- DIFFA: Large Language Diffusion Models Can Listen and UnderstandJiaming Zhou, Hongjie Chen, Shiwan Zhao, Jian Kang et al.AAAI 2026 · 10 citations
- Closing the Modality Reasoning Gap for Speech Large Language ModelsChaoren Wang, Heng Lu, Xueyao Zhang, Shujie Liu et al.ACL 2026 · 10 citations
Builds on14
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- SALMONN: Towards Generic Hearing Abilities for Large Language ModelsChangli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen et al.ICLR 2024 · 557 citations
- Listen, Think, and UnderstandYuan Gong, Hongyin Luo, Alexander H. Liu, Leonid Karlinsky et al.ICLR 2024 · 247 citations
- SLURP: A Spoken Language Understanding Resource PackageEmanuele Bastianelli, Andrea Vanzo, Pawel Swietojanski, Verena RieserEMNLP 2020 · 129 citations
- InSerter: Speech Instruction Following with Unsupervised Interleaved Pre-trainingDingdong Wang, Jin Xu, Ruihang Chu, Zhifang Guo et al.ACL 2025 · 9 citations
Related papers
- HPSU: A Benchmark for Human-Level Perception in Real-World Spoken Speech UnderstandingChen Li, Peiji Yang, Yicheng Zhong, Jianxing Yu et al.AAAI 2026 · 1 citation
- HumanSense: From Multimodal Perception to Empathetic Context-Aware Responses Through Reasoning MLLMsZheng Qin, Ruobing Zheng, Yabing Wang, Tianqi Li et al.AAAI 2026 · 2 citations
- HAVE-Bench: Hierarchical Audio-Visual Evaluation from Perception to InteractionZhong Muyan, Erfei Cui, Sen Xing, Weiyun Wang et al.CVPR 2026
- MULTIVOX: A Benchmark for Evaluating Voice Assistants for Multimodal InteractionsRamaneswaran Selvakumar, Ashish Seth, Nishit Anand, Utkarsh Tyagi et al.EMNLP 2025
- Beyond Emotion Recognition: A Multi-Turn Multimodal Emotion Understanding and Reasoning BenchmarkJinpeng Hu, Hongchang Shi, Chongyuan Dai, Zhuo Li et al.ACM MM 2025 · 7 citations
