CoHear: Conversation Enhancement via Multi-earphone Collaboration
Lixing He, Yunqi Guo, Zhenyu Yan, Guoliang Xing
Abstract
In crowded social settings like conferences, background noise, overlapping voices, and lively interactions often lead to "cocktail party deafness, " hindering clear conversation. While modern earphones are a promising platform for speech enhancement, existing solutions are limited: they either operate on a single device, ignoring the multi-party nature of conversation, or rely on impractical assumptions like fixed conversation areas and pre-recorded audio. We present CoHear, a collaborative system that leverages a network of earphones to holistically model and enhance speech at the conversation level. CoHear bridges acoustic sensor networks with deep learning for target speech extraction through two key contributions: 1) a novel, conversation-driven network that dynamically forms groups based on user interaction, using verbal and non-verbal cues (primarily head orientation) for robust, infrastructure-free coordination; and 2) a bandwidth-efficient, robust target speech extraction model that effectively utilizes peer-relayed audio as conditioning signals, even under network constraints. CoHear is evaluated in both real-world experiments and simulations. Results show that our conversation network obtains more than 90% accuracy in group formation, improves the speech quality by up to 8.8 dB over state-of-the-art baselines, and demonstrates real-time performance on a mobile device. In a user study with 20 participants, CoHear has a much higher score than baseline with good usability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 47c070d4-13bb-47ef-b311-f4f10f2bca0fBuilds on20
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis et al.CVPR 2022 · 525 citations
- Push the Limit of Acoustic Gesture RecognitionYanwen Wang, Jiaxing Shen, Yuanqing ZhengINFOCOM 2020 · 80 citations
- The Cone of Silence: Speech Separation by LocalizationTeerapat Jenrungrot, Vivek Jayaram, Steven M. Seitz, Ira Kemelmacher-ShlizermanNeurIPS 2020 · 70 citations
- UltraSE: single-channel speech enhancement using ultrasoundKe Sun, Xinyu ZhangMobiCom 2021 · 69 citations
- Ear-AR: indoor acoustic augmented reality on earphonesZhijian Yang, Yu-Lin Wei, Sheng Shen, Romit Roy ChoudhuryMobiCom 2020 · 47 citations
Related papers
- ClearSpeech: Improving Voice Quality of Earbuds Using Both In-Ear and Out-Ear MicrophonesDong Ma, Ting Dang, Ming Ding, Rajesh BalanUbiComp 2024 · 5 citations
- EarSE: Bringing Robust Speech Enhancement to COTS HeadphonesDi Duan, Yongliang Chen, Weitao Xu, Tianxing LiUbiComp 2024 · 14 citations
- EarSpeech: Exploring In-Ear Occlusion Effect on Earphones for Data-efficient Airborne Speech EnhancementFeiyu Han, Panlong Yang, You Zuo, Fei Shang et al.UbiComp 2024 · 10 citations
- Egocentric Deep Multi-Channel Audio-Visual Active Speaker LocalizationHao Jiang, Calvin Murdock, Vamsi Krishna IthapuCVPR 2022 · 39 citations
- DeepEar: Sound Localization with Binaural MicrophonesQiang Yang, Yuanqing ZhengINFOCOM 2022 · 8 citations
