CLARIS: Clear and Intelligible Speech from Whispered and Dysarthric Voices
Neil Shah, Yash Sonkar, Shirish Subhash Karande, Vineet Gandhi
摘要
Whispered and dysarthric speech hinder effective communication and undermine the reliability of voice-enabled systems. We present CLARIS, a compact speech-to-speech restoration system that turns such atypical input into clear, expressive speech. CLARIS requires no disorder-specific architectural tuning, generalizes across languages, and adapts quickly to new accents and speakers, enabling practical personalization. On whispered English, Hindi, and clinically challenging dysarthric speech, CLARIS delivers state-of-the-art intelligibility and naturalness, with listener studies confirming gains in quality, intelligibility, naturalness, and prosody. The system runs in real time, converting one second of input in about 30ms and enables inclusive, private, and personalized voice interaction. Audio samples are available at https://claris-w2s.github.io/CLARIS/
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- WESPER: Zero-shot and Realtime Whisper to Normal Voice Conversion for Whisper-based Speech InteractionsJun RekimotoCHI 2023 · 被引用 24 次
- Idiosyncratic Versus Normative Modeling of Atypical Speech Recognition: Dysarthric Case StudiesVishnu Raja, Adithya V. Ganesan, Anand Syamkumar, Ritwik Banerjee 等EMNLP 2025 · 被引用 2 次
- EA-VAE: Learning to Reconstruct Dysarthric Speech via Variational Autoencoder with Encoding AlignmentDaipeng Zhang, Wenhuan Lu, Xianghu Yue, Hongcheng Zhang 等AAAI 2026
- DualVoice: Speech Interaction that Discriminates between Normal and Whispered Voice InputJun RekimotoUIST 2022 · 被引用 9 次
- From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech RecognitionTianduo Wang, Lu Xu, Wei Lu, Shanbo ChengEMNLP 2025 · 被引用 1 次
