USpeech: Ultrasound-Enhanced Speech with Minimal Human Effort via Cross-Modal Synthesis
Luca Jiang-Tao Yu, Running Zhao, Sijie Ji, Edith C. H. Ngai, Chenshu Wu
Abstract
Speech enhancement is crucial for ubiquitous human-computer interaction. Recently, ultrasound-based acoustic sensing has emerged as an attractive choice for speech enhancement because of its superior ubiquity and performance. However, due to inevitable interference from unexpected and unintended sources during audio-ultrasound data acquisition, existing solutions rely heavily on human effort for data collection and processing. This leads to significant data scarcity that limits the full potential of ultrasound-based speech enhancement. To address this, we propose USpeech, a cross-modal ultrasound synthesis framework for speech enhancement with minimal human effort. At its core is a two-stage framework that establishes the correspondence between visual and ultrasonic modalities by leveraging audio as a bridge. This approach overcomes challenges from the lack of paired video-ultrasound datasets and the inherent heterogeneity between video and ultrasound data. Our framework incorporates contrastive video-audio pre-training to project modalities into a shared semantic space and employs an audio-ultrasound encoder-decoder for ultrasound synthesis. We then present a speech enhancement network that enhances speech in the time-frequency domain and recovers the clean speech waveform via a neural vocoder. Comprehensive experiments show USpeech achieves remarkable performance using synthetic ultrasound data comparable to physical data, outperforming state-of-the-art ultrasound-based speech enhancement baselines. USpeech is open-sourced at https://github.com/aiot-lab/USpeech/.
CCS Concepts: • Human-centered computing → Ubiquitous and mobile computing design and evaluation methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on23
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Learning Audio-Visual Speech Representation by Masked Multimodal Cluster PredictionBowen Shi, Wei-Ning Hsu, Kushal Lakhotia, Abdelrahman MohamedICLR 2022 · 460 citations
- PHASEN: A Phase-and-Harmonics-Aware Speech Enhancement NetworkDacheng Yin, Chong Luo, Zhiwei Xiong, Wenjun ZengAAAI 2020 · 387 citations
- IMUTube: Automatic Extraction of Virtual on-body Accelerometry from Video for Human Activity RecognitionHyeokHyen Kwon, Catherine Tong, Harish Haresamudram, Yan Gao et al.UbiComp 2020 · 153 citations
Related papers
- UltraSpeech: Speech Enhancement by Interaction between Ultrasound and SpeechHan Ding, Yizhan Wang, Hao Li, Cui Zhao et al.UbiComp 2022 · 32 citations
- UltraSE: single-channel speech enhancement using ultrasoundKe Sun, Xinyu ZhangMobiCom 2021 · 69 citations
- Sensing to Hear through Memory: Ultrasound Speech Enhancement without Real Ultrasound SignalsQian Zhang, Ke Liu, Dong WangUbiComp 2024 · 5 citations
- Multichannel AV-wav2vec2: A Framework for Learning Multichannel Multi-Modal Speech RepresentationQiushi Zhu, Jie Zhang, Yu Gu, Yuchen Hu et al.AAAI 2024 · 17 citations
- CLAPSpeech: Learning Prosody from Text Context with Contrastive Language-Audio Pre-TrainingZhenhui Ye, Rongjie Huang, Yi Ren, Ziyue Jiang et al.ACL 2023 · 13 citations
