WESPER: Zero-shot and Realtime Whisper to Normal Voice Conversion for Whisper-based Speech Interactions
Jun Rekimoto
摘要
Recognizing whispered speech and converting it to normal speech creates many possibilities for speech interaction. Because the sound pressure of whispered speech is significantly lower than that of normal speech, it can be used as a semi-silent speech interaction in public places without being audible to others. Converting whispers to normal speech also improves the speech quality for people with speech or hearing impairments. However, conventional speech conversion techniques do not provide sufficient conversion quality or require speaker-dependent datasets consisting of pairs of whispered and normal speech utterances. To address these problems, we propose WESPER, a zero-shot, real-time whisper-to-normal speech conversion mechanism based on self-supervised learning. WESPER consists of a speech-to-unit (STU) encoder, which generates hidden speech units common to both whispered and normal speech, and a unit-to-speech (UTS) decoder, which reconstructs speech from the encoded speech units. Unlike the existing methods, this conversion is user-independent and does not require a paired dataset for whispered and normal speech. The UTS decoder can reconstruct speech in any target speaker’s voice from speech units, and it requires only an unlabeled target speaker’s speech data. We confirmed that the quality of the speech converted from a whisper was improved while preserving its natural prosody. Additionally, we confirmed the effectiveness of the proposed approach to perform speech reconstruction for people with speech or hearing disabilities.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Sensible Agent: A Framework for Unobtrusive Interaction with Proactive AR AgentsGeonsun Lee, Min Xia, Nels Numan, Xun Qian 等UIST 2025 · 被引用 15 次
- Mouse2Vec: Learning Reusable Semantic Representations of Mouse BehaviourGuanhua Zhang, Zhiming Hu, Mihai Bâce, Andreas BullingCHI 2024 · 被引用 5 次
- NasoVoce: A Nose-Mounted Low-Audibility Speech Interface for Always-Available Speech InteractionJun Rekimoto, Yu Nishimura, Bojian YangCHI 2026 · 被引用 1 次
- CRAFT: Exploring Wearable Creative AI on Smart Glasses for Fiction Writing in Real-World ContextsRunze Cai, Yuxuan Huang, Lin-Ping Yuan, Kexin Xiang 等UbiComp 2026
它引用的顶会 Paper3
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 被引用 2,890 次
- FastSpeech 2: Fast and High-Quality End-to-End Text to SpeechYi Ren, Chenxu Hu, Xu Tan, Tao Qin 等ICLR 2021 · 被引用 513 次
- DualVoice: Speech Interaction that Discriminates between Normal and Whispered Voice InputJun RekimotoUIST 2022 · 被引用 9 次
相关 Paper
- Self-Taught Recognizer: Toward Unsupervised Adaptation for Speech Foundation ModelsYuchen Hu, Chen Chen, Chao-Han Huck Yang, Chengwei Qin 等NeurIPS 2024 · 被引用 14 次
- CLARIS: Clear and Intelligible Speech from Whispered and Dysarthric VoicesNeil Shah, Yash Sonkar, Shirish Subhash Karande, Vineet GandhiCHI 2026 · 被引用 1 次
- Unsupervised Speech RecognitionAlexei Baevski, Wei-Ning Hsu, Alexis Conneau, Michael AuliNeurIPS 2021 · 被引用 309 次
- StethoSpeech: Speech Generation Through a Clinical Stethoscope Attached to the SkinNeil Kumar Shah, Neha Sahipjohn, Vishal Tambrahalli, Ramanathan Subramanian 等UbiComp 2024 · 被引用 5 次
- Self-supervised learning with random-projection quantizer for speech recognitionChung-Cheng Chiu, James Qin, Yu Zhang, Jiahui Yu 等ICML 2022 · 被引用 245 次
