EA-VAE: Learning to Reconstruct Dysarthric Speech via Variational Autoencoder with Encoding Alignment
Daipeng Zhang, Wenhuan Lu, Xianghu Yue, Hongcheng Zhang, Jianguo Wei
Abstract
Dysarthric speech reconstruction (DSR) aims to enhance the intelligibility of dysarthric speech. Compared with normal speech, the dysarthric speech is characterized by its pathological features, including discontinuous pronunciation, slow speech, hoarseness, and improper pauses. Significant disparities in the feature space between normal and dysarthric speech may result in suboptimal speech reconstruction, thereby degrading speech intelligibility. To enhance the reconstruction ability of speech feature spaces, this paper proposes a DSR model named the Encoding-Aligned Variational Autoencoder (EA-VAE). By incorporating alignment modules of frame-level embedding features, prior distributions, and duration into the encoder of the VAE, the model explicitly aligns the dysarthric speech encoding with a representation of the parallel normal speech. A shared decoder is then used to generate speech with improved intelligibility. Experimental results on the UASpeech benchmark confirm that EA-VAE achieves state-of-the-art performance, with a 31.7% relative word error rate reduction and the highest subjective MOS score (4.48), thoroughly validating the effectiveness and advancements of the proposed method in dysarthric speech reconstruction.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e0d111bb-0b11-4ea2-a2e8-624d3ffc95b6Builds on3
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 2,890 citations
- Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-SpeechJaehyeon Kim, Jungil Kong, Juhee SonICML 2021 · 1,267 citations
- Glow-TTS: A Generative Flow for Text-to-Speech via Monotonic Alignment SearchJaehyeon Kim, Sungwon Kim, Jungil Kong, Sungroh YoonNeurIPS 2020 · 663 citations
Related papers
- Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech SynthesisWeiwei Lin, Chenhang HeICLR 2025
- Cross-Utterance Conditioned VAE for Non-Autoregressive Text-to-SpeechYang Li, Cheng Yu, Guangzhi Sun, Hua Jiang et al.ACL 2022 · 7 citations
- Revisit and Outstrip Entity Alignment: A Perspective of Generative ModelsLingbing Guo, Zhuo Chen, Jiaoyan Chen, Yin Fang et al.ICLR 2024 · 17 citations
- A Variational Framework for Improving Naturalness in Generative Spoken Language ModelsLi-Wei Chen, Takuya Higuchi, Zakaria Aldeneh, Ahmed Hussen Abdelaziz et al.ICML 2025
- Is Discourse Role Important for Emotion Recognition in Conversation?Donovan Ong, Jian Su, Bin Chen, Anh Tuan Luu et al.AAAI 2022 · 29 citations
