Hybrid Neural Networks for On-Device Directional Hearing
Anran Wang, Maruchi Kim, Hao Zhang, Shyamnath Gollakota
Abstract
On-device directional hearing requires audio source separation from a given direction while achieving stringent human-imperceptible latency requirements. While neural nets can achieve significantly better performance than traditional beamformers, all existing models fall short of supporting low-latency causal inference on computationally-constrained wearables. We present Hybrid-Beam, a hybrid model that combines traditional beamformers with a custom lightweight neural net. The former reduces the computational burden of the latter and also improves its generalizability, while the latter is designed to further reduce the memory and computational overhead to enable real-time and low-latency operations. Our evaluation shows comparable performance to state-of-the-art causal inference models on synthetic data while achieving a 5x reduction of model size, 4x reduction of computation per second, 5x reduction in processing time and generalizing better to real hardware data. Further, our real-time hybrid model runs in 8 ms on mobile CPUs designed for low-power wearable devices and achieves an end-to-end latency of 17.5 ms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 25a14573-afce-48cb-89ae-e27e78f7ad71Cited by top-tier papers5
- SLNet: A Spectrogram Learning Neural Network for Deep Wireless SensingZheng Yang, Yi Zhang, Kun Qian, Chenshu WuNSDI 2023 · 57 citations
- Semantic Hearing: Programming Acoustic Scenes with Binaural HearablesBandhav Veluri, Malek Itani, Justin Chan, Takuya Yoshioka et al.UIST 2023 · 29 citations
- Look Once to Hear: Target Speech Hearing with Noisy ExamplesBandhav Veluri, Malek Itani, Tuochao Chen, Takuya Yoshioka et al.CHI 2024 · 25 citations
- Spatial Speech Translation: Translating Across Space With Binaural HearablesTuochao Chen, Qirui Wang, Runlin He, Shyamnath GollakotaCHI 2025 · 5 citations
- SonicSieve: Bringing Directional Speech Extraction to Smartphones Using Acoustic MicrostructuresKuang Yuan, Yifeng Wang, Xiyuxing Zhang, Chengyi Shen et al.CHI 2026 · 1 citation
Builds on2
Related papers
- Wireless Hearables With Programmable Speech AI AcceleratorsMalek Itani, Tuochao Chen, Arun Raghavan, Gavriel Kohlberg et al.MobiCom 2025 · 4 citations
- Hypoformer: Hybrid Decomposition Transformer for Edge-friendly Neural Machine TranslationSunzhu Li, Peng Zhang, Guobing Gan, Xiuqing Lv et al.EMNLP 2022 · 3 citations
- TRAMBA: A Hybrid Transformer and Mamba Architecture for Practical Audio and Bone Conduction Speech Super Resolution and Enhancement on Mobile and Wearable PlatformsYueyuan Sui, Minghui Zhao, Junxi Xia, Xiaofan Jiang et al.UbiComp 2025 · 18 citations
- MobilePortrait: Real-Time One-Shot Neural Head Avatars on Mobile DevicesJianwen Jiang, Gaojie Lin, Zhengkun Rong, Chao Liang et al.CVPR 2025
- Short-Term Memory ConvolutionsGrzegorz Stefanski, Krzysztof Arendt, Pawel Daniluk, Bartlomiej Jasik et al.ICLR 2023 · 67 citations
