QualiSpeech: A Speech Quality Assessment Dataset with Natural Language Reasoning and Descriptions
Siyin Wang, Wenyi Yu, Xianzhao Chen, Xiaohai Tian, Jun Zhang, Lu Lu, Yu Tsao, Junichi Yamagishi, Yuxuan Wang, Chao Zhang
摘要
This paper explores a novel perspective to speech quality assessment by leveraging natural language descriptions, offering richer, more nuanced insights than traditional numerical scoring methods. Natural language feedback provides instructive recommendations and detailed evaluations, yet existing datasets lack the comprehensive annotations needed for this approach. To bridge this gap, we introduce QualiSpeech, a comprehensive low-level speech quality assessment dataset encompassing 11 key aspects and detailed natural language comments that include reasoning and contextual insights. Additionally, we propose the QualiSpeech Benchmark to evaluate the low-level speech understanding capabilities of auditory large language models (LLMs). Experimental results demonstrate that finetuned auditory LLMs can reliably generate detailed descriptions of noise and distortion, effectively identifying their types and temporal characteristics. The results further highlight the potential for incorporating reasoning to enhance the accuracy and reliability of quality assessments. The dataset will be released at https://huggingface.co/datasets/tsinghua-ee/QualiSpeech.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- SpeechJudge: Towards Human-Level Judgment for Speech NaturalnessXueyao Zhang, Chaoren Wang, Huan Liao, Ziniu Li 等ICLR 2026 · 被引用 32 次
- SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality EvaluationHui Wang, Jinghua Zhao, Yifan Yang, Shujie Liu 等ACL 2026 · 被引用 21 次
- ARECHO: Autoregressive Evaluation via Chain-Based Hypothesis Optimization for Speech Multi-Metric EstimationJiatong Shi, Yifan Cheng, Bo-Hao Su, Hye-jin Shim 等NeurIPS 2025 · 被引用 4 次
它引用的顶会 Paper9
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- BEATs: Audio Pre-Training with Acoustic TokenizersSanyuan Chen, Yu Wu, Chengyi Wang, Shujie Liu 等ICML 2023 · 被引用 568 次
- SECap: Speech Emotion Captioning with Large Language ModelYaoxun Xu, Hangting Chen, Jianwei Yu, Qiaochu Huang 等AAAI 2024 · 被引用 70 次
- Audio Entailment: Assessing Deductive Reasoning for Audio UnderstandingSoham Deshmukh, Shuo Han, Hazim T. Bukhari, Benjamin Elizalde 等AAAI 2025 · 被引用 23 次
- SpeechCraft: A Fine-Grained Expressive Speech Dataset with Natural Language DescriptionZeyu Jin, Jia Jia, Qixin Wang, Kehan Li 等ACM MM 2024 · 被引用 12 次
相关 Paper
- Audio Large Language Models Can Be Descriptive Speech Quality EvaluatorsChen Chen, Yuchen Hu, Siyin Wang, Helin Wang 等ICLR 2025
- HPSU: A Benchmark for Human-Level Perception in Real-World Spoken Speech UnderstandingChen Li, Peiji Yang, Yicheng Zhong, Jianxing Yu 等AAAI 2026 · 被引用 1 次
- Read to Hear: A Zero-Shot Pronunciation Assessment Using Textual Descriptions and LLMsYu-Wen Chen, Melody Ma, Julia HirschbergEMNLP 2025
- Towards Holistic Evaluation of Large Audio-Language Models: A Comprehensive SurveyChih-Kai Yang, Neo S. Ho, Hung-yi LeeEMNLP 2025 · 被引用 7 次
- SpeechIQ: Speech-Agentic Intelligence Quotient Across Cognitive Levels in Voice Understanding by Large Language ModelsZhen Wan, Chao-Han Huck Yang, Yahan Yu, Jinchuan Tian 等ACL 2025 · 被引用 2 次
