APG-MOS: Auditory Perception Guided-MOS Predictor for Synthetic Speech
Zhicheng Lian, Lizhi Wang, Hua Huang
摘要
Automatic speech quality assessment aims to quantify subjective human perception of speech through computational models to reduce the need for labor-consuming manual evaluations. While models based on deep learning have achieved progress in predicting mean opinion scores (MOS) to assess synthetic speech, the neglect of fundamental auditory perception mechanisms limits consistency with human judgments. To address this issue, we propose an auditory perception guided-MOS prediction model (APG-MOS) that synergistically integrates auditory modeling with semantic analysis to enhance consistency with human judgments. Specifically, we first design a perceptual module, grounded in biological auditory mechanisms, to simulate cochlear functions, which encodes acoustic signals into biologically aligned electrochemical representations. Secondly, we propose a residual vector quantization (RVQ)-based semantic distortion modeling method to quantify the degradation of speech quality at the semantic level. Finally, we design a residual cross-attention architecture, coupled with a progressive learning strategy, to enable multimodal fusion of encoded electrochemical signals and semantic representations. Experiments demonstrate that APG-MOS achieves superior performance on two primary benchmarks. Our code and checkpoint will be available on a public repository upon publication.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper18
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- SALMONN: Towards Generic Hearing Abilities for Large Language ModelsChangli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen 等ICLR 2024 · 被引用 557 次
- ProDiff: Progressive Fast Diffusion Model for High-Quality Text-to-SpeechRongjie Huang, Zhou Zhao, Huadai Liu, Jinglin Liu 等ACM MM 2022 · 被引用 182 次
- SCOREQ: Speech Quality Assessment with Contrastive RegressionAlessandro Ragano, Jan Skoglund, Andrew HinesNeurIPS 2024 · 被引用 90 次
- NORESQA: A Framework for Speech Quality Assessment using Non-Matching ReferencesPranay Manocha, Buye Xu, Anurag KumarNeurIPS 2021 · 被引用 67 次
相关 Paper
- Audio Large Language Models Can Be Descriptive Speech Quality EvaluatorsChen Chen, Yuchen Hu, Siyin Wang, Helin Wang 等ICLR 2025
- ADGNet: Attention Discrepancy Guided Deep Neural Network for Blind Image Quality AssessmentXiaoyu Ma, Yaqi Wang, Chang Liu, Suiyu Zhang 等ACM MM 2022 · 被引用 6 次
- Self-Supervised Speech Quality Estimation and Enhancement Using Only Clean SpeechSzu-Wei Fu, Kuo-Hsuan Hung, Yu Tsao, Yu-Chiang Frank WangICLR 2024 · 被引用 27 次
- DR.Experts: Differential Refinement of Distortion-Aware Experts for Blind Image Quality AssessmentBohan Fu, Guanyi Qin, Fazhan Zhang, Zihao Huang 等AAAI 2026 · 被引用 1 次
- SemGes: Semantics-Aware Co-Speech Gesture Generation Using Semantic Coherence and Relevance LearningLanmiao Liu, Esam Ghaleb, Asli Özyürek, Zerrin YumakICCV 2025 · 被引用 4 次
