StethoSpeech: Speech Generation Through a Clinical Stethoscope Attached to the Skin
Neil Kumar Shah, Neha Sahipjohn, Vishal Tambrahalli, Ramanathan Subramanian, Vineet Gandhi
Abstract
We introduce StethoSpeech, a silent speech interface that transforms flesh-conducted vibrations behind the ear into speech. This innovation is designed to improve social interactions for those with voice disorders, and furthermore enable discreet public communication. Unlike prior efforts, StethoSpeech does not require (a) paired-speech data for recorded vibrations and (b) a specialized device for recording vibrations, as it can work with an off-the-shelf clinical stethoscope. The novelty of our framework lies in the overall design, simulation of the ground-truth speech, and a sequence-to-sequence translation network, which works in the latent space. We present comprehensive experiments on the existing CSTR NAM TIMIT Plus corpus and our proposed StethoText: a large-scale synchronized database of non-audible murmur and text for speech research. Our results show that StethoSpeech provides natural-sounding and intelligible speech, significantly outperforming existing methods on several quantitative and qualitative metrics. Additionally, we showcase its capacity to extend its application to speakers not encountered during training and its effectiveness in challenging, noisy environments. Speech samples are available at https://stethospeech.github.io/StethoSpeech/.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 7ddec45f-914c-4ce4-a143-5b8ec0a5ca9dRelated papers
- EchoSpeech: Continuous Silent Speech Recognition on Minimally-obtrusive Eyewear Powered by Acoustic SensingRuidong Zhang, Ke Li, Yihong Hao, Yufan Wang et al.CHI 2023 · 53 citations
- WESPER: Zero-shot and Realtime Whisper to Normal Voice Conversion for Whisper-based Speech InteractionsJun RekimotoCHI 2023 · 24 citations
- ReHEarSSE: Recognizing Hidden-in-the-Ear Silently Spelled ExpressionsXuefu Dong, Yifei Chen, Yuuki Nishiyama, Kaoru Sezaki et al.CHI 2024 · 20 citations
- NasoVoce: A Nose-Mounted Low-Audibility Speech Interface for Always-Available Speech InteractionJun Rekimoto, Yu Nishimura, Bojian YangCHI 2026 · 1 citation
- From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech RecognitionTianduo Wang, Lu Xu, Wei Lu, Shanbo ChengEMNLP 2025 · 1 citation
