Self-Supervised Speech Quality Estimation and Enhancement Using Only Clean Speech
Szu-Wei Fu, Kuo-Hsuan Hung, Yu Tsao, Yu-Chiang Frank Wang
Abstract
Speech quality estimation has recently undergone a paradigm shift from human-hearing expert designs to machine-learning models. However, current models rely mainly on supervised learning, which is time-consuming and expensive for label collection. To solve this problem, we propose VQScore, a self-supervised metric for evaluating speech based on the quantization error of a vector-quantized-variational autoencoder (VQ-VAE). The training of VQ-VAE relies on clean speech; hence, large quantization errors can be expected when the speech is distorted. To further improve correlation with real quality scores, domain knowledge of speech processing is incorporated into the model design. We found that the vector quantization mechanism could also be used for self-supervised speech enhancement (SE) model training. To improve the robustness of the encoder for SE, a novel self-distillation mechanism combined with adversarial training is introduced. In summary, the proposed speech quality estimation method and enhancement models require only clean speech for training without any label requirements. Experimental results show that the proposed VQScore and enhancement model are competitive with supervised baselines. The code will be released after publication.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- SCOREQ: Speech Quality Assessment with Contrastive RegressionAlessandro Ragano, Jan Skoglund, Andrew HinesNeurIPS 2024 · 90 citations
- MAPSS: Manifold-based Assessment of Perceptual Source SeparationAmir Ivry, Samuele Cornell, Shinji WatanabeICLR 2026 · 1 citation
- GenSE: Generative Speech Enhancement via Language Models using Hierarchical ModelingJixun Yao, Hexin Liu, Chen Chen, Yuchen Hu et al.ICLR 2025
- GOMPSNR: Reflourish the Signal-to-Noise Ratio Metric for Audio Generation TasksLingling Dai, Andong Li, Cheng Chi, Yifan Liang et al.AAAI 2026
Builds on5
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self DistillationLinfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen et al.ICCV 2019 · 1,069 citations
- Vector-quantized Image Modeling with Improved VQGANJiahui Yu, Xin Li, Jing Yu Koh, Han Zhang et al.ICLR 2022 · 753 citations
- Self-supervised learning with random-projection quantizer for speech recognitionChung-Cheng Chiu, James Qin, Yu Zhang, Jiahui Yu et al.ICML 2022 · 245 citations
- NORESQA: A Framework for Speech Quality Assessment using Non-Matching ReferencesPranay Manocha, Buye Xu, Anurag KumarNeurIPS 2021 · 67 citations
Related papers
- SQ-VAE: Variational Bayes on Discrete Representation with Self-annealed Stochastic QuantizationYuhta Takida, Takashi Shibuya, Wei-Hsiang Liao, Chieh-Hsin Lai et al.ICML 2022 · 99 citations
- Self-Guidance: Enhancing Neural Codecs via Decoder Manifold AlignmentXiang Li, Yixuan Zhou, Xie, Zhiyong Wu et al.ICML 2026
- SECodec: Structural Entropy-based Compressive Speech Representation Codec for Speech Language ModelsLinqin Wang, Yaping Liu, Zhengtao Yu, Shengxiang Gao et al.AAAI 2025 · 3 citations
- DiVeQ: Differentiable Vector Quantization Using the Reparameterization TrickMohammad Hassan Vali, Tom Bäckström, Arno SolinICLR 2026 · 6 citations
- Ada-DQA: Adaptive Diverse Quality-aware Feature Acquisition for Video Quality AssessmentHongbo Liu, Mingda Wu, Kun Yuan, Ming Sun et al.ACM MM 2023 · 18 citations
