EarSpeech: Exploring In-Ear Occlusion Effect on Earphones for Data-efficient Airborne Speech Enhancement
Feiyu Han, Panlong Yang, You Zuo, Fei Shang, Fenglei Xu, Xiang-Yang Li
Abstract
Earphones have become a popular voice input and interaction device. However, airborne speech is susceptible to ambient noise, making it necessary to improve the quality and intelligibility of speech on earphones in noisy conditions. As the dual-microphone structure (i.e., outer and in-ear microphones) has been widely adopted in earphones (especially ANC earphones), we design EarSpeech which exploits in-ear acoustic sensory as the complementary modality to enable airborne speech enhancement. The key idea of EarSpeech is that in-ear speech is less sensitive to ambient noise and exhibits a correlation with airborne speech. However, due to the occlusion effect, in-ear speech has limited bandwidth, making it challenging to directly correlate with full-band airborne speech. Therefore, we exploit the occlusion effect to carry out theoretical modeling and quantitative analysis of this cross-channel correlation and study how to leverage such cross-channel correlation for speech enhancement. Specifically, we design a series of methodologies including data augmentation, deep learning-based fusion, and noise mixture scheme, to improve the generalization, effectiveness, and robustness of EarSpeech, respectively. Lastly, we conduct real-world experiments to evaluate the performance of our system. Specifically, EarSpeech achieves an average improvement ratio of 27.23% and 13.92% in terms of PESQ and STOI, respectively, and significantly improves SI-SDR by 8.91 dB. Benefiting from data augmentation, EarSpeech can achieve comparable performance with a small-scale dataset that is 40 times less than the original dataset. In addition, we validate the generalization of different users, speech content, and language types, respectively, as well as robustness in the real world via comprehensive experiments. The audio demo of EarSpeech is available on https://github.com/EarSpeech/earspeech.github.io/.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers3
- A Survey of Earable Technology: Trends, Tools, and the Road AheadChangshuo Hu, Qiang Yang, Yang Liu, Tobias Röddiger et al.UbiComp 2026 · 4 citations
- FeelWave: Enabling Emotion-Aware Voice Interaction through Noise-Robust mmWave Emotion SensingLingyu Wang, You Zuo, Dequan Wang, Chenming He et al.CHI 2026 · 1 citation
- CoHear: Conversation Enhancement via Multi-earphone CollaborationLixing He, Yunqi Guo, Zhenyu Yan, Guoliang XingUbiComp 2026 · 1 citation
Related papers
- Exploring and Addressing Low-Quality Auxiliary Modality in Earable Dual-microphone Speech EnhancementFeiyu Han, Dawei Yan, Shanyue Wang, Jinyang Huang et al.UbiComp 2026
- ClearSpeech: Improving Voice Quality of Earbuds Using Both In-Ear and Out-Ear MicrophonesDong Ma, Ting Dang, Ming Ding, Rajesh BalanUbiComp 2024 · 5 citations
- EarAcE: Empowering Versatile Acoustic Sensing via Earable Active Noise Cancellation PlatformYetong Cao, Chao Cai, Anbo Yu, Fan Li et al.UbiComp 2023 · 22 citations
- UltraSpeech: Speech Enhancement by Interaction between Ultrasound and SpeechHan Ding, Yizhan Wang, Hao Li, Cui Zhao et al.UbiComp 2022 · 32 citations
- EarSE: Bringing Robust Speech Enhancement to COTS HeadphonesDi Duan, Yongliang Chen, Weitao Xu, Tianxing LiUbiComp 2024 · 14 citations
