APG-MOS: Auditory Perception Guided-MOS Predictor for Synthetic Speech
Zhicheng Lian, Lizhi Wang, Hua Huang
Abstract
Automatic speech quality assessment aims to quantify subjective human perception of speech through computational models to reduce the need for labor-consuming manual evaluations. While models based on deep learning have achieved progress in predicting mean opinion scores (MOS) to assess synthetic speech, the neglect of fundamental auditory perception mechanisms limits consistency with human judgments. To address this issue, we propose an auditory perception guided-MOS prediction model (APG-MOS) that synergistically integrates auditory modeling with semantic analysis to enhance consistency with human judgments. Specifically, we first design a perceptual module, grounded in biological auditory mechanisms, to simulate cochlear functions, which encodes acoustic signals into biologically aligned electrochemical representations. Secondly, we propose a residual vector quantization (RVQ)-based semantic distortion modeling method to quantify the degradation of speech quality at the semantic level. Finally, we design a residual cross-attention architecture, coupled with a progressive learning strategy, to enable multimodal fusion of encoded electrochemical signals and semantic representations. Experiments demonstrate that APG-MOS achieves superior performance on two primary benchmarks. Our code and checkpoint will be available on a public repository upon publication.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 225b1d40-42ee-4148-9623-97eceec5ffe3Cited by top-tier papers1
Ask how each one uses itBuilds on18
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- SALMONN: Towards Generic Hearing Abilities for Large Language ModelsChangli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen et al.ICLR 2024 · 557 citations
- ProDiff: Progressive Fast Diffusion Model for High-Quality Text-to-SpeechRongjie Huang, Zhou Zhao, Huadai Liu, Jinglin Liu et al.ACM MM 2022 · 182 citations
- SCOREQ: Speech Quality Assessment with Contrastive RegressionAlessandro Ragano, Jan Skoglund, Andrew HinesNeurIPS 2024 · 90 citations
- NORESQA: A Framework for Speech Quality Assessment using Non-Matching ReferencesPranay Manocha, Buye Xu, Anurag KumarNeurIPS 2021 · 67 citations
Related papers
- Audio Large Language Models Can Be Descriptive Speech Quality EvaluatorsChen Chen, Yuchen Hu, Siyin Wang, Helin Wang et al.ICLR 2025
- ADGNet: Attention Discrepancy Guided Deep Neural Network for Blind Image Quality AssessmentXiaoyu Ma, Yaqi Wang, Chang Liu, Suiyu Zhang et al.ACM MM 2022 · 6 citations
- Self-Supervised Speech Quality Estimation and Enhancement Using Only Clean SpeechSzu-Wei Fu, Kuo-Hsuan Hung, Yu Tsao, Yu-Chiang Frank WangICLR 2024 · 27 citations
- DR.Experts: Differential Refinement of Distortion-Aware Experts for Blind Image Quality AssessmentBohan Fu, Guanyi Qin, Fazhan Zhang, Zihao Huang et al.AAAI 2026 · 1 citation
- SemGes: Semantics-Aware Co-Speech Gesture Generation Using Semantic Coherence and Relevance LearningLanmiao Liu, Esam Ghaleb, Asli Özyürek, Zerrin YumakICCV 2025 · 4 citations
