CoHear: Conversation Enhancement via Multi-earphone Collaboration
Lixing He, Yunqi Guo, Zhenyu Yan, Guoliang Xing
摘要
In crowded social settings like conferences, background noise, overlapping voices, and lively interactions often lead to "cocktail party deafness, " hindering clear conversation. While modern earphones are a promising platform for speech enhancement, existing solutions are limited: they either operate on a single device, ignoring the multi-party nature of conversation, or rely on impractical assumptions like fixed conversation areas and pre-recorded audio. We present CoHear, a collaborative system that leverages a network of earphones to holistically model and enhance speech at the conversation level. CoHear bridges acoustic sensor networks with deep learning for target speech extraction through two key contributions: 1) a novel, conversation-driven network that dynamically forms groups based on user interaction, using verbal and non-verbal cues (primarily head orientation) for robust, infrastructure-free coordination; and 2) a bandwidth-efficient, robust target speech extraction model that effectively utilizes peer-relayed audio as conditioning signals, even under network constraints. CoHear is evaluated in both real-world experiments and simulations. Results show that our conversation network obtains more than 90% accuracy in group formation, improves the speech quality by up to 8.8 dB over state-of-the-art baselines, and demonstrates real-time performance on a mobile device. In a user study with 20 participants, CoHear has a much higher score than baseline with good usability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper20
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis 等CVPR 2022 · 被引用 525 次
- Push the Limit of Acoustic Gesture RecognitionYanwen Wang, Jiaxing Shen, Yuanqing ZhengINFOCOM 2020 · 被引用 80 次
- The Cone of Silence: Speech Separation by LocalizationTeerapat Jenrungrot, Vivek Jayaram, Steven M. Seitz, Ira Kemelmacher-ShlizermanNeurIPS 2020 · 被引用 70 次
- UltraSE: single-channel speech enhancement using ultrasoundKe Sun, Xinyu ZhangMobiCom 2021 · 被引用 69 次
- Ear-AR: indoor acoustic augmented reality on earphonesZhijian Yang, Yu-Lin Wei, Sheng Shen, Romit Roy ChoudhuryMobiCom 2020 · 被引用 47 次
相关 Paper
- ClearSpeech: Improving Voice Quality of Earbuds Using Both In-Ear and Out-Ear MicrophonesDong Ma, Ting Dang, Ming Ding, Rajesh BalanUbiComp 2024 · 被引用 5 次
- EarSE: Bringing Robust Speech Enhancement to COTS HeadphonesDi Duan, Yongliang Chen, Weitao Xu, Tianxing LiUbiComp 2024 · 被引用 14 次
- EarSpeech: Exploring In-Ear Occlusion Effect on Earphones for Data-efficient Airborne Speech EnhancementFeiyu Han, Panlong Yang, You Zuo, Fei Shang 等UbiComp 2024 · 被引用 10 次
- Egocentric Deep Multi-Channel Audio-Visual Active Speaker LocalizationHao Jiang, Calvin Murdock, Vamsi Krishna IthapuCVPR 2022 · 被引用 39 次
- DeepEar: Sound Localization with Binaural MicrophonesQiang Yang, Yuanqing ZhengINFOCOM 2022 · 被引用 8 次
