UltraSE: single-channel speech enhancement using ultrasound
Ke Sun, Xinyu Zhang
Abstract
Robust speech enhancement is considered as the holy grail of audio processing and a key requirement for human-human and human-machine interaction. Solving this task with single-channel, audio-only methods remains an open challenge, especially for practical scenarios involving a mixture of competing speakers and background noise. In this paper, we propose UltraSE, which uses ultrasound sensing as a complementary modality to separate the desired speaker's voice from interferences and noise. UltraSE uses a commodity mobile device (e.g., smartphone) to emit ultrasound and capture the reflections from the speaker's articulatory gestures. It introduces a multi-modal, multi-domain deep learning framework to fuse the ultrasonic Doppler features and the audible speech spectrogram. Furthermore, it employs an adversarially trained discriminator, based on a cross-modal similarity measurement network, to learn the correlation between the two heterogeneous feature modalities. Our experiments verify that UltraSE simultaneously improves speech intelligibility and quality, and outperforms state-of-the-art solutions by a large margin.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 862ecda2-f97a-45e6-a1cb-3f3bf4142a91Cited by top-tier papers14
- mmMIC: Multi-modal Speech Recognition based on mmWave RadarLong Fan, Lei Xie, Xinran Lu, Yi Li et al.INFOCOM 2023 · 41 citations
- PowerPhone: Unleashing the Acoustic Sensing Capability of SmartphonesShirui Cao, Dong Li, Sunghoon Ivan Lee, Jie XiongMobiCom 2023 · 35 citations
- Exploring the Feasibility of Remote Cardiac Auscultation Using EarphonesTao Chen, Yongjie Yang, Xiaoran Fan, Xiuzhen Guo et al.MobiCom 2024 · 34 citations
- mSilent: Towards General Corpus Silent Speech Recognition Using COTS mmWave RadarShang Zeng, Haoran Wan, Shuyu Shi, Wei WangUbiComp 2023 · 34 citations
- InertiEAR: Automatic and Device-independent IMU-based Eavesdropping on SmartphonesMing Gao, Yajie Liu, Yike Chen, Yimin Li et al.INFOCOM 2022 · 18 citations
Builds on5
- PHASEN: A Phase-and-Harmonics-Aware Speech Enhancement NetworkDacheng Yin, Chong Luo, Zhiwei Xiong, Wenjun ZengAAAI 2020 · 387 citations
- Hearing Your Voice is Not Enough: An Articulatory Gesture Based Liveness Detection for Voice AuthenticationLinghan Zhang, Sheng Tan, Jie YangCCS 2017 · 212 citations
- Using Sonar for Liveness Detection to Protect Smart Speakers against Remote AttackersYeonjoon Lee, Yue Zhao, Jiutian Zeng, Kwangwuk Lee et al.UbiComp 2020 · 36 citations
- VocalLock: Sensing Vocal Tract for Passphrase-Independent User Authentication Leveraging Acoustic Signals on SmartphonesLi Lu, Jiadi Yu, Yingying Chen, Yan WangUbiComp 2020 · 35 citations
- Dynamic Speed Warping: Similarity-Based One-shot Learning for Device-free Gesture SignalsXun Wang, Ke Sun, Ting Zhao, Wei Wang et al.INFOCOM 2020 · 21 citations
Related papers
- UltraSpeech: Speech Enhancement by Interaction between Ultrasound and SpeechHan Ding, Yizhan Wang, Hao Li, Cui Zhao et al.UbiComp 2022 · 32 citations
- EarSE: Bringing Robust Speech Enhancement to COTS HeadphonesDi Duan, Yongliang Chen, Weitao Xu, Tianxing LiUbiComp 2024 · 14 citations
- Sensing to Hear through Memory: Ultrasound Speech Enhancement without Real Ultrasound SignalsQian Zhang, Ke Liu, Dong WangUbiComp 2024 · 5 citations
- USpeech: Ultrasound-Enhanced Speech with Minimal Human Effort via Cross-Modal SynthesisLuca Jiang-Tao Yu, Running Zhao, Sijie Ji, Edith C. H. Ngai et al.UbiComp 2025 · 1 citation
- mmWave-Aided Unified Speech Enhancement and Separation without Speaker Count PriorDachao Han, Teng Huang, Han Ding, Cui Zhao et al.INFOCOM 2026 · 1 citation
