Hybrid Neural Networks for On-Device Directional Hearing
Anran Wang, Maruchi Kim, Hao Zhang, Shyamnath Gollakota
摘要
On-device directional hearing requires audio source separation from a given direction while achieving stringent human-imperceptible latency requirements. While neural nets can achieve significantly better performance than traditional beamformers, all existing models fall short of supporting low-latency causal inference on computationally-constrained wearables. We present Hybrid-Beam, a hybrid model that combines traditional beamformers with a custom lightweight neural net. The former reduces the computational burden of the latter and also improves its generalizability, while the latter is designed to further reduce the memory and computational overhead to enable real-time and low-latency operations. Our evaluation shows comparable performance to state-of-the-art causal inference models on synthetic data while achieving a 5x reduction of model size, 4x reduction of computation per second, 5x reduction in processing time and generalizing better to real hardware data. Further, our real-time hybrid model runs in 8 ms on mobile CPUs designed for low-power wearable devices and achieves an end-to-end latency of 17.5 ms.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- SLNet: A Spectrogram Learning Neural Network for Deep Wireless SensingZheng Yang, Yi Zhang, Kun Qian, Chenshu WuNSDI 2023 · 被引用 57 次
- Semantic Hearing: Programming Acoustic Scenes with Binaural HearablesBandhav Veluri, Malek Itani, Justin Chan, Takuya Yoshioka 等UIST 2023 · 被引用 29 次
- Look Once to Hear: Target Speech Hearing with Noisy ExamplesBandhav Veluri, Malek Itani, Tuochao Chen, Takuya Yoshioka 等CHI 2024 · 被引用 25 次
- Spatial Speech Translation: Translating Across Space With Binaural HearablesTuochao Chen, Qirui Wang, Runlin He, Shyamnath GollakotaCHI 2025 · 被引用 5 次
- SonicSieve: Bringing Directional Speech Extraction to Smartphones Using Acoustic MicrostructuresKuang Yuan, Yifeng Wang, Xiyuxing Zhang, Chengyi Shen 等CHI 2026 · 被引用 1 次
它引用的顶会 Paper2
相关 Paper
- Wireless Hearables With Programmable Speech AI AcceleratorsMalek Itani, Tuochao Chen, Arun Raghavan, Gavriel Kohlberg 等MobiCom 2025 · 被引用 4 次
- Hypoformer: Hybrid Decomposition Transformer for Edge-friendly Neural Machine TranslationSunzhu Li, Peng Zhang, Guobing Gan, Xiuqing Lv 等EMNLP 2022 · 被引用 3 次
- TRAMBA: A Hybrid Transformer and Mamba Architecture for Practical Audio and Bone Conduction Speech Super Resolution and Enhancement on Mobile and Wearable PlatformsYueyuan Sui, Minghui Zhao, Junxi Xia, Xiaofan Jiang 等UbiComp 2025 · 被引用 18 次
- MobilePortrait: Real-Time One-Shot Neural Head Avatars on Mobile DevicesJianwen Jiang, Gaojie Lin, Zhengkun Rong, Chao Liang 等CVPR 2025
- Short-Term Memory ConvolutionsGrzegorz Stefanski, Krzysztof Arendt, Pawel Daniluk, Bartlomiej Jasik 等ICLR 2023 · 被引用 67 次
