UltraSE: single-channel speech enhancement using ultrasound
Ke Sun, Xinyu Zhang
摘要
Robust speech enhancement is considered as the holy grail of audio processing and a key requirement for human-human and human-machine interaction. Solving this task with single-channel, audio-only methods remains an open challenge, especially for practical scenarios involving a mixture of competing speakers and background noise. In this paper, we propose UltraSE, which uses ultrasound sensing as a complementary modality to separate the desired speaker's voice from interferences and noise. UltraSE uses a commodity mobile device (e.g., smartphone) to emit ultrasound and capture the reflections from the speaker's articulatory gestures. It introduces a multi-modal, multi-domain deep learning framework to fuse the ultrasonic Doppler features and the audible speech spectrogram. Furthermore, it employs an adversarially trained discriminator, based on a cross-modal similarity measurement network, to learn the correlation between the two heterogeneous feature modalities. Our experiments verify that UltraSE simultaneously improves speech intelligibility and quality, and outperforms state-of-the-art solutions by a large margin.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- mmMIC: Multi-modal Speech Recognition based on mmWave RadarLong Fan, Lei Xie, Xinran Lu, Yi Li 等INFOCOM 2023 · 被引用 41 次
- PowerPhone: Unleashing the Acoustic Sensing Capability of SmartphonesShirui Cao, Dong Li, Sunghoon Ivan Lee, Jie XiongMobiCom 2023 · 被引用 35 次
- Exploring the Feasibility of Remote Cardiac Auscultation Using EarphonesTao Chen, Yongjie Yang, Xiaoran Fan, Xiuzhen Guo 等MobiCom 2024 · 被引用 34 次
- mSilent: Towards General Corpus Silent Speech Recognition Using COTS mmWave RadarShang Zeng, Haoran Wan, Shuyu Shi, Wei WangUbiComp 2023 · 被引用 34 次
- InertiEAR: Automatic and Device-independent IMU-based Eavesdropping on SmartphonesMing Gao, Yajie Liu, Yike Chen, Yimin Li 等INFOCOM 2022 · 被引用 18 次
它引用的顶会 Paper5
- PHASEN: A Phase-and-Harmonics-Aware Speech Enhancement NetworkDacheng Yin, Chong Luo, Zhiwei Xiong, Wenjun ZengAAAI 2020 · 被引用 387 次
- Hearing Your Voice is Not Enough: An Articulatory Gesture Based Liveness Detection for Voice AuthenticationLinghan Zhang, Sheng Tan, Jie YangCCS 2017 · 被引用 212 次
- Using Sonar for Liveness Detection to Protect Smart Speakers against Remote AttackersYeonjoon Lee, Yue Zhao, Jiutian Zeng, Kwangwuk Lee 等UbiComp 2020 · 被引用 36 次
- VocalLock: Sensing Vocal Tract for Passphrase-Independent User Authentication Leveraging Acoustic Signals on SmartphonesLi Lu, Jiadi Yu, Yingying Chen, Yan WangUbiComp 2020 · 被引用 35 次
- Dynamic Speed Warping: Similarity-Based One-shot Learning for Device-free Gesture SignalsXun Wang, Ke Sun, Ting Zhao, Wei Wang 等INFOCOM 2020 · 被引用 21 次
相关 Paper
- UltraSpeech: Speech Enhancement by Interaction between Ultrasound and SpeechHan Ding, Yizhan Wang, Hao Li, Cui Zhao 等UbiComp 2022 · 被引用 32 次
- EarSE: Bringing Robust Speech Enhancement to COTS HeadphonesDi Duan, Yongliang Chen, Weitao Xu, Tianxing LiUbiComp 2024 · 被引用 14 次
- Sensing to Hear through Memory: Ultrasound Speech Enhancement without Real Ultrasound SignalsQian Zhang, Ke Liu, Dong WangUbiComp 2024 · 被引用 5 次
- USpeech: Ultrasound-Enhanced Speech with Minimal Human Effort via Cross-Modal SynthesisLuca Jiang-Tao Yu, Running Zhao, Sijie Ji, Edith C. H. Ngai 等UbiComp 2025 · 被引用 1 次
- mmWave-Aided Unified Speech Enhancement and Separation without Speaker Count PriorDachao Han, Teng Huang, Han Ding, Cui Zhao 等INFOCOM 2026 · 被引用 1 次
