PASE: Leveraging the Phonological Prior of WavLM for Low-Hallucination Generative Speech Enhancement
Xiaobin Rong, Qinwen Hu, Mansur Yesilbursa, Kamil Wójcicki, Jing Lu
Abstract
Generative models have shown remarkable performance in speech enhancement (SE), achieving superior perceptual quality over traditional discriminative approaches. However, existing generative SE approaches often overlook the risk of hallucination under severe noise, leading to incorrect spoken content or inconsistent speaker characteristics, which we term linguistic and acoustic hallucinations, respectively. We argue that linguistic hallucination stems from models' failure to constrain valid phonological structures and it is a more fundamental challenge. While language models (LMs) are well-suited for capturing the underlying speech structure through modeling the distribution of discrete tokens, existing approaches are limited in learning from noise-corrupted representations, which can lead to contaminated priors and hallucinations. To overcome these limitations, we propose the Phonologically Anchored Speech Enhancer (PASE), a generative SE framework that leverages the robust phonological prior embedded in the pre-trained WavLM model to mitigate hallucinations. First, we adapt WavLM into a denoising expert via representation distillation to clean its final-layer features. Guided by the model's intrinsic phonological prior, this process enables robust denoising while minimizing linguistic hallucinations. To further reduce acoustic hallucinations, we train the vocoder with a dual-stream representation: the high-level phonetic representation provides clean linguistic content, while a low-level acoustic representation retains speaker identity and prosody. Experimental results demonstrate that PASE not only surpasses state-of-the-art discriminative models in perceptual quality, but also significantly outperforms prior generative models with substantially lower linguistic and acoustic hallucinations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 86f22a89-5b11-4a16-97ca-59176de6a640Builds on7
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 2,890 citations
- High-Fidelity Audio Compression with Improved RVQGANRithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar et al.NeurIPS 2023 · 910 citations
- Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesisHubert SiuzdakICLR 2024 · 229 citations
- GenSE: Generative Speech Enhancement via Language Models using Hierarchical ModelingJixun Yao, Hexin Liu, Chen Chen, Yuchen Hu et al.ICLR 2025
Related papers
- LLaSE-G1: Incentivizing Generalization Capability for LLaMA-based Speech EnhancementBoyi Kang, Xinfa Zhu, Zihan Zhang, Zhen Ye et al.ACL 2025
- From Continuous to Discrete: Cross-Domain Collaborative General Speech Enhancement via Hierarchical Language ModelsZhaoxi Mu, Rilin Chen, Andong Li, Meng Yu et al.ACM MM 2025
- Hierarchical Semantic-Acoustic Modeling via Semi-Discrete Residual Representations for Expressive End-to-End Speech SynthesisYixuan Zhou, Guoyang Zeng, Xin Liu, Xiang Li et al.ICLR 2026
- FINALLY: fast and universal speech enhancement with studio-like qualityNicholas Babaev, Kirill Tamogashev, Azat Saginbaev, Ivan Shchekotov et al.NeurIPS 2024 · 28 citations
- Listen like a Teacher: Mitigating Whisper Hallucinations Using Adaptive Layer Attention and Knowledge DistillationKumud Tripathi, Aditya Srinivas Menon, Aman Gaurav, Raj Prakash Gohil et al.AAAI 2026
