WearSE: Enabling Streaming Speech Enhancement on Eyewear Using Acoustic Sensing
Qian Zhang, Kaiyi Guo, Yifei Yang, Dong Wang
Abstract
Smart eyewear has rapidly evolved in recent years, yet its mobile and in-the-wild characteristics often make voice interactions on such devices susceptible to external interferences. In this paper, we introduce WearSE, a system that utilizes acoustic signals emitted and received by speakers and microphones mounted on eyewear to perceive facial movements during speech, achieving multimodal speech enhancement. WearSE incorporates three key designs to meet the high demands for real-time operation and robustness on smart eyewear. First, considering the frequent use in mobile scenarios, we design a sensing-enhanced network to amplify the capability of acoustic sensing, eliminating dynamic multipath interferences. Second, we develop a lightweight speech enhancement network that enhances both the amplitude and phase of the speech spectrum. Through a casual network design, computational demands are significantly reduced, ensuring real-time operation on mobile devices. Third, addressing the scarcity of paired data, we design a memory-based back-translation mechanism to generate pseudo-acoustic sensing data using a large amount of publicly available speech data for network training. We construct a prototype system and extensively evaluate WearSE through experiments. In multi-speaker scenarios, our approach exhibits much better performance than pure audio speech enhancement methods. Comparisons with commercial smart eyewear also demonstrate that WearSE significantly surpasses existing noise reduction algorithms in these devices. The audio demo of WearSE is available on https://github.com/WearSE/wearse.github.io.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get effdb009-19f8-4624-b095-e8c521561206Cited by top-tier papers1
Ask how each one uses itRelated papers
- mmMUSE: An mmWave-based Motion-resilient Universal Speech Enhancement SystemLingyu Wang, Kai Wang, Dequan Wang, You Zuo et al.UbiComp 2026 · 2 citations
- Acoustic-based Upper Facial Action Recognition for Smart EyewearWentao Xie, Qian Zhang, Jin ZhangUbiComp 2021 · 31 citations
- EarSE: Bringing Robust Speech Enhancement to COTS HeadphonesDi Duan, Yongliang Chen, Weitao Xu, Tianxing LiUbiComp 2024 · 14 citations
- EyeEcho: Continuous and Low-power Facial Expression Tracking on GlassesKe Li, Ruidong Zhang, Siyuan Chen, Boao Chen et al.CHI 2024 · 27 citations
- Sensing to Hear: Speech Enhancement for Mobile Devices Using Acoustic SignalsQian Zhang, Dong Wang, Run Zhao, Yinggang Yu et al.UbiComp 2021 · 22 citations
