WESPER: Zero-shot and Realtime Whisper to Normal Voice Conversion for Whisper-based Speech Interactions
Jun Rekimoto
Abstract
Recognizing whispered speech and converting it to normal speech creates many possibilities for speech interaction. Because the sound pressure of whispered speech is significantly lower than that of normal speech, it can be used as a semi-silent speech interaction in public places without being audible to others. Converting whispers to normal speech also improves the speech quality for people with speech or hearing impairments. However, conventional speech conversion techniques do not provide sufficient conversion quality or require speaker-dependent datasets consisting of pairs of whispered and normal speech utterances. To address these problems, we propose WESPER, a zero-shot, real-time whisper-to-normal speech conversion mechanism based on self-supervised learning. WESPER consists of a speech-to-unit (STU) encoder, which generates hidden speech units common to both whispered and normal speech, and a unit-to-speech (UTS) decoder, which reconstructs speech from the encoded speech units. Unlike the existing methods, this conversion is user-independent and does not require a paired dataset for whispered and normal speech. The UTS decoder can reconstruct speech in any target speaker’s voice from speech units, and it requires only an unlabeled target speaker’s speech data. We confirmed that the quality of the speech converted from a whisper was improved while preserving its natural prosody. Additionally, we confirmed the effectiveness of the proposed approach to perform speech reconstruction for people with speech or hearing disabilities.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6dbee1fc-3ff3-4d48-9c66-a105d0498224Cited by top-tier papers4
- Sensible Agent: A Framework for Unobtrusive Interaction with Proactive AR AgentsGeonsun Lee, Min Xia, Nels Numan, Xun Qian et al.UIST 2025 · 15 citations
- Mouse2Vec: Learning Reusable Semantic Representations of Mouse BehaviourGuanhua Zhang, Zhiming Hu, Mihai Bâce, Andreas BullingCHI 2024 · 5 citations
- NasoVoce: A Nose-Mounted Low-Audibility Speech Interface for Always-Available Speech InteractionJun Rekimoto, Yu Nishimura, Bojian YangCHI 2026 · 1 citation
- CRAFT: Exploring Wearable Creative AI on Smart Glasses for Fiction Writing in Real-World ContextsRunze Cai, Yuxuan Huang, Lin-Ping Yuan, Kexin Xiang et al.UbiComp 2026
Builds on3
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 2,890 citations
- FastSpeech 2: Fast and High-Quality End-to-End Text to SpeechYi Ren, Chenxu Hu, Xu Tan, Tao Qin et al.ICLR 2021 · 513 citations
- DualVoice: Speech Interaction that Discriminates between Normal and Whispered Voice InputJun RekimotoUIST 2022 · 9 citations
Related papers
- Self-Taught Recognizer: Toward Unsupervised Adaptation for Speech Foundation ModelsYuchen Hu, Chen Chen, Chao-Han Huck Yang, Chengwei Qin et al.NeurIPS 2024 · 14 citations
- CLARIS: Clear and Intelligible Speech from Whispered and Dysarthric VoicesNeil Shah, Yash Sonkar, Shirish Subhash Karande, Vineet GandhiCHI 2026 · 1 citation
- Unsupervised Speech RecognitionAlexei Baevski, Wei-Ning Hsu, Alexis Conneau, Michael AuliNeurIPS 2021 · 309 citations
- StethoSpeech: Speech Generation Through a Clinical Stethoscope Attached to the SkinNeil Kumar Shah, Neha Sahipjohn, Vishal Tambrahalli, Ramanathan Subramanian et al.UbiComp 2024 · 5 citations
- Self-supervised learning with random-projection quantizer for speech recognitionChung-Cheng Chiu, James Qin, Yu Zhang, Jiahui Yu et al.ICML 2022 · 245 citations
