DeepEar: Sound Localization with Binaural Microphones
Qiang Yang, Yuanqing Zheng
Abstract
Binaural microphones, referring to two microphones with artificial human-shaped ears, are pervasively used in humanoid robots and hearing aids improving sound quality. In many applications, it is crucial for such robots to interact with humans by finding the voice direction. However, sound source localization with binaural microphones remains challenging, especially in multi-source scenarios. Prior works utilize microphone arrays to deal with the multi-source localization problem. Extra arrays yet incur higher deployment costs and take up more space. However, human brains have evolved to locate multiple sound sources with only two ears. Inspired by this fact, we propose DeepEar, a binaural microphone-based localization system that can locate multiple sounds. To this end, we develop a neural network to mimic the acoustic signal processing pipeline of the human auditory system. Different from hand-crafted features used in prior works, DeepEar can automatically extract useful features for localization. More importantly, the trained neural networks can be extended and adapted to new environments with a minimum amount of extra training data. Experiment results show that DeepEar can substantially outperform the state-of-the-art deep learning approach, with a sound detection accuracy of 93.3% and an azimuth estimation error of 7.4 degrees in multisource scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- RFBoost: Understanding and Boosting Deep WiFi Sensing via Physical Data AugmentationWeiying Hou, Chenshu WuUbiComp 2024 · 24 citations
- AutoIOT: LLM-Driven Automated Natural Language Programming for AIoT ApplicationsLeming Shen, Qiang Yang, Yuanqing Zheng, Mo LiMobiCom 2025 · 14 citations
- VoShield: Voice Liveness Detection with Sound Field DynamicsQiang Yang, Kaiyan Cui, Yuanqing ZhengINFOCOM 2023 · 13 citations
- A Survey of Earable Technology: Trends, Tools, and the Road AheadChangshuo Hu, Qiang Yang, Yang Liu, Tobias Röddiger et al.UbiComp 2026 · 4 citations
- ISDrama: Immersive Spatial Drama Generation through Multimodal PromptingYu Zhang, Wenxiang Guo, Changhao Pan, Zhiyuan Zhu et al.ACM MM 2025 · 1 citation
Builds on3
- Voice localization using nearby wall reflectionsSheng Shen, Daguan Chen, Yu-Lin Wei, Zhijian Yang et al.MobiCom 2020 · 79 citations
- Personalizing head related transfer functions for earablesZhijian Yang, Romit Roy ChoudhurySIGCOMM 2021 · 26 citations
- AcouRadar: Towards Single Source based Acoustic LocalizationLinsong Cheng, Zhao Wang, Yunting Zhang, Weiyi Wang et al.INFOCOM 2020 · 13 citations
Related papers
- Learning to Separate Voices by Spatial RegionsAlan Xu, Romit Roy ChoudhuryICML 2022 · 18 citations
- Binaural Audio-Visual LocalizationXinyi Wu, Zhenyao Wu, Lili Ju, Song WangAAAI 2021 · 32 citations
- CoHear: Conversation Enhancement via Multi-earphone CollaborationLixing He, Yunqi Guo, Zhenyu Yan, Guoliang XingUbiComp 2026 · 1 citation
- SonicSieve: Bringing Directional Speech Extraction to Smartphones Using Acoustic MicrostructuresKuang Yuan, Yifeng Wang, Xiyuxing Zhang, Chengyi Shen et al.CHI 2026 · 1 citation
- Neural Synthesis of Binaural Speech From Mono AudioAlexander Richard, Dejan Markovic, Israel D. Gebru, Steven Krenn et al.ICLR 2021 · 73 citations
