Audio Large Language Models Can Be Descriptive Speech Quality Evaluators
Chen Chen, Yuchen Hu, Siyin Wang, Helin Wang, Zhehuai Chen, Chao Zhang, Chao-Han Huck Yang, Eng Siong Chng
摘要
An ideal multimodal agent should be aware of the quality of its input modalities. Recent advances have enabled large language models (LLMs) to incorporate auditory systems for handling various speech-related tasks. However, most audio LLMs remain unaware of the quality of the speech they process. This limitation arises because speech quality evaluation is typically excluded from multi-task training due to the lack of suitable datasets. To address this, we introduce the first natural language-based speech evaluation corpus, generated from authentic human ratings. In addition to the overall Mean Opinion Score (MOS), this corpus offers detailed analysis across multiple dimensions and identifies causes of quality degradation. It also enables descriptive comparisons between two speech samples (A/B tests) with human-like judgment. Leveraging this corpus, we propose an alignment approach with LLM distillation (ALLD) to guide the audio LLM in extracting relevant information from raw speech and generating meaningful responses. Experimental results demonstrate that ALLD outperforms the previous state-of-the-art regression model in MOS prediction, with a mean square error of 0.17 and an A/B test accuracy of 98.6%. Additionally, the generated responses achieve BLEU scores of 25.8 and 30.2 on two tasks, surpassing the capabilities of task-specific models. This work advances the comprehensive perception of speech signals by audio LLMs, contributing to the development of real-world auditory and sensory intelligent agents.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- SpeechJudge: Towards Human-Level Judgment for Speech NaturalnessXueyao Zhang, Chaoren Wang, Huan Liao, Ziniu Li 等ICLR 2026 · 被引用 32 次
- SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality EvaluationHui Wang, Jinghua Zhao, Yifan Yang, Shujie Liu 等ACL 2026 · 被引用 21 次
- QualiSpeech: A Speech Quality Assessment Dataset with Natural Language Reasoning and DescriptionsSiyin Wang, Wenyi Yu, Xianzhao Chen, Xiaohai Tian 等ACL 2025 · 被引用 20 次
- Flex-Judge: Text-Only Reasoning Unleashes Zero-Shot Multimodal EvaluatorsJongwoo Ko, Sungnyun Kim, Sungwoo Cho, Se-Young YunNeurIPS 2025 · 被引用 6 次
- ARECHO: Autoregressive Evaluation via Chain-Based Hypothesis Optimization for Speech Multi-Metric EstimationJiatong Shi, Yifan Cheng, Bo-Hao Su, Hye-jin Shim 等NeurIPS 2025 · 被引用 4 次
它引用的顶会 Paper11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context LearningHaokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta 等NeurIPS 2022 · 被引用 1,483 次
相关 Paper
- Closing the Gap Between Text and Speech Understanding in LLMsSantiago Cuervo, Skyler Seto, Maureen de Seyssel, Richard He Bai 等ICLR 2026 · 被引用 18 次
- APG-MOS: Auditory Perception Guided-MOS Predictor for Synthetic SpeechZhicheng Lian, Lizhi Wang, Hua HuangACM MM 2025 · 被引用 1 次
- Understanding the Modality Gap: An Empirical Study on the Speech-Text Alignment Mechanism of Large Speech Language ModelsBajian Xiang, Shuaijiang Zhao, Tingwei Guo, Wei ZouEMNLP 2025 · 被引用 6 次
- Making Visual Dialogue More Engaging: A New Task, Method, and MetricGuanghui Ye, Huan Zhao, Yingxue Gao, Zhixue Zhao 等AAAI 2026
- Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-ProbabilityYong Ren, Jingbei Li, Haiyang Sun, Yujie Chen 等ICML 2026
