EarSE: Bringing Robust Speech Enhancement to COTS Headphones
Di Duan, Yongliang Chen, Weitao Xu, Tianxing Li
Abstract
Speech enhancement is regarded as the key to the quality of digital communication and is gaining increasing attention in the research field of audio processing. In this paper, we present EarSE, the first robust, hands-free, multi-modal speech enhancement solution using commercial off-the-shelf headphones. The key idea of EarSE is a novel hardware setting---leveraging the form factor of headphones equipped with a boom microphone to establish a stable acoustic sensing field across the user's face. Furthermore, we designed a sensing methodology based on Frequency-Modulated Continuous-Wave, which is an ultrasonic modality sensitive to capture subtle facial articulatory gestures of users when speaking. Moreover, we design a fully attention-based deep neural network to self-adaptively solve the user diversity problem by introducing the Vision Transformer network. We enhance the collaboration between the speech and ultrasonic modalities using a multi-head attention mechanism and a Factorized Bilinear Pooling gate. Extensive experiments demonstrate that EarSE achieves remarkable performance as increasing SiSDR by 14.61 dB and reducing the word error rate of user speech recognition by 22.45--66.41% in real-world application. EarSE not only outperforms seven baselines by 38.0% in SiSNR, 12.4% in STOI, and 20.5% in PESQ on average but also maintains practicality.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get eb8bad7f-0d54-4d30-ae59-04a470c519a2Cited by top-tier papers4
- TRAMBA: A Hybrid Transformer and Mamba Architecture for Practical Audio and Bone Conduction Speech Super Resolution and Enhancement on Mobile and Wearable PlatformsYueyuan Sui, Minghui Zhao, Junxi Xia, Xiaofan Jiang et al.UbiComp 2025 · 18 citations
- A Survey of Earable Technology: Trends, Tools, and the Road AheadChangshuo Hu, Qiang Yang, Yang Liu, Tobias Röddiger et al.UbiComp 2026 · 4 citations
- CoHear: Conversation Enhancement via Multi-earphone CollaborationLixing He, Yunqi Guo, Zhenyu Yan, Guoliang XingUbiComp 2026 · 1 citation
- USpeech: Ultrasound-Enhanced Speech with Minimal Human Effort via Cross-Modal SynthesisLuca Jiang-Tao Yu, Running Zhao, Sijie Ji, Edith C. H. Ngai et al.UbiComp 2025 · 1 citation
Related papers
- UltraSE: single-channel speech enhancement using ultrasoundKe Sun, Xinyu ZhangMobiCom 2021 · 69 citations
- WearSE: Enabling Streaming Speech Enhancement on Eyewear Using Acoustic SensingQian Zhang, Kaiyi Guo, Yifei Yang, Dong WangUbiComp 2025 · 7 citations
- UltraSpeech: Speech Enhancement by Interaction between Ultrasound and SpeechHan Ding, Yizhan Wang, Hao Li, Cui Zhao et al.UbiComp 2022 · 32 citations
- ClearSpeech: Improving Voice Quality of Earbuds Using Both In-Ear and Out-Ear MicrophonesDong Ma, Ting Dang, Ming Ding, Rajesh BalanUbiComp 2024 · 5 citations
- ReHEarSSE: Recognizing Hidden-in-the-Ear Silently Spelled ExpressionsXuefu Dong, Yifei Chen, Yuuki Nishiyama, Kaoru Sezaki et al.CHI 2024 · 20 citations
