emg2speech: synthesizing speech from electromyography using self-supervised speech models
Harshavardhana T. Gowda, Daniel C. Comstock, Lee M. Miller
摘要
We present a neuromuscular speech interface that translates electromyographic (EMG) signals recorded from orofacial muscles during speech articulation directly into audio. We find that self-supervised speech (S3) representations are strongly linearly related to the electrical power of muscle activity: a simple linear mapping predicts EMG power from S3 representations with a correlation of r = 0.85. In addition, EMG power vectors associated with distinct articulatory gestures form structured, separable clusters. Together, these observations suggest that S3 models implicitly encode articulatory mechanisms, as reflected in EMG activity. Leveraging this structure, we map EMG signals into the S3 representation space and synthesize speech, enabling end-to-end EMG-to-speech generation without explicit articulatory modeling or vocoder training. We demonstrate this system with a participant with amyotrophic lateral sclerosis (ALS), converting orofacial EMG recorded while she silently articulated speech into audio. PROJECT PAGE. GITHUB. DATA.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- Lip to Speech Synthesis with Visual Context Attentional GANMinsu Kim, Joanna Hong, Yong Man RoNeurIPS 2021 · 被引用 76 次
- Digital Voicing of Silent SpeechDavid Gaddy, Dan KleinEMNLP 2020 · 被引用 55 次
- Plug-and-Play Stability for Intracortical Brain-Computer Interfaces: A One-Year Demonstration of Seamless Brain-to-Text CommunicationChaofei Fan, Nick Hahn, Foram Kamdar, Donald T. Avansino 等NeurIPS 2023 · 被引用 36 次
相关 Paper
- Reading Your Actions: Learning Generalizable Action Representations via Pre-training AEMGZhenghao Huang, Huilin Yao, Kaikai Wang, Lin ShuCVPR 2026
- Translating Signals to Languages for sEMG-Based Activity RecognitionMing Wang, Haoxuan Qu, Qiuhong Ke, Wei Zhou 等CVPR 2026
- Interpretable Self-Supervised Facial Micro-Expression Learning to Predict Cognitive State and Neurological DisordersArun Das, Jeffrey Mock, Yufei Huang, Edward J. Golob 等AAAI 2021 · 被引用 11 次
- Towards Accurate Lip-to-Speech Synthesis in-the-WildSindhu B. Hegde, Rudrabha Mukhopadhyay, C. V. Jawahar, Vinay P. NamboodiriACM MM 2023 · 被引用 9 次
- SLD-L2S: Hierarchical Subspace Latent Diffusion for High-Fidelity Lip to Speech SynthesisYifan Liang, Andong Li, Kang Yang, Guochen Yu 等AAAI 2026
