MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks
Yadong Niu, TIANZI WANG, Heinrich Dinkel, Xingwei Sun, Jiahao Zhou, Gang Li, Jizhong Liu, Xunying Liu, Jian Luan
Abstract
While large audio-language models have advanced open-ended audio understanding, they still fall short of nuanced human-level comprehension. This gap persists largely because current benchmarks, limited by data annotations and evaluation metrics, fail to reliably distinguish between generic and highly detailed model outputs. To this end, this work introduces MECAT, a Multi-Expert Constructed Benchmark for Fine-Grained Audio Understanding Tasks. Generated via a pipeline that integrates analysis from specialized expert models with Chain-of-Thought large language model reasoning, MECAT provides multi-perspective, fine-grained captions and open-set question-answering pairs. The benchmark is complemented by a novel metric: DATE (Discriminative-Enhanced Audio Text Evaluation). This metric penalizes generic terms and rewards detailed descriptions by combining single-sample semantic similarity with crosssample discriminability. A comprehensive evaluation of state-of-the-art audio models is also presented, providing new insights into their current capabilities and limitations. The code and data are publicly available at https://github.com/ xiaomi-research/mecat .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on10
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- SALMONN: Towards Generic Hearing Abilities for Large Language ModelsChangli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen et al.ICLR 2024 · 557 citations
- Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language ModelsSreyan Ghosh, Arushi Goel, Jaehyeon Kim, Sonal Kumar et al.NeurIPS 2025 · 299 citations
- MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised TrainingYizhi Li, Ruibin Yuan, Ge Zhang, Yinghao Ma et al.ICLR 2024 · 277 citations
- Learning to Answer Questions in Dynamic Audio-Visual ScenariosGuangyao Li, Yake Wei, Yapeng Tian, Chenliang Xu et al.CVPR 2022 · 101 citations
Related papers
- AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video GenerationZiwei Zhou, Zeyuan Lai, Rui Wang, Yifan Yang et al.ICML 2026 · 8 citations
- Fine-grained Audible Video DescriptionXuyang Shen, Dong Li, Jinxing Zhou, Zhen Qin et al.CVPR 2023
- STAR-Bench: Probing Deep Spatio-Temporal Reasoning as Audio 4D IntelligenceZihan Liu, Zhikang Niu, Qiuyang Xiao, Zhisheng Zheng et al.ICLR 2026 · 12 citations
- HalluAudio: A Comprehensive Benchmark for Hallucination Detection in Large Audio-Language ModelsFeiyu Zhao, Yiming Chen, Wenhuan Lu, Daipeng Zhang et al.ACL 2026 · 3 citations
- MMAU: A Massive Multi-Task Audio Understanding and Reasoning BenchmarkS. Sakshi, Utkarsh Tyagi, Sonal Kumar, Ashish Seth et al.ICLR 2025
