Semantic Hearing: Programming Acoustic Scenes with Binaural Hearables
Bandhav Veluri, Malek Itani, Justin Chan, Takuya Yoshioka, Shyamnath Gollakota
摘要
Imagine being able to listen to the birds chirping in a park without hearing the chatter from other hikers, or being able to block out traffic noise on a busy street while still being able to hear emergency sirens and car honks. We introduce semantic hearing, a novel capability for hearable devices that enables them to, in real-time, focus on, or ignore, specific sounds from real-world environments, while also preserving the spatial cues. To achieve this, we make two technical contributions: 1) we present the first neural network that can achieve binaural target sound extraction in the presence of interfering sounds and background noise, and 2) we design a training methodology that allows our system to generalize to real-world use. Results show that our system can operate with 20 sound classes and that our transformer-based network has a runtime of 6.56 ms on a connected smartphone. In-the-wild evaluation with participants in previously unseen indoor and outdoor scenarios shows that our proof-of-concept system can extract the target sounds and generalize to preserve the spatial cues in its binaural output. Project page with code: https://semantichearing.cs.washington.edu
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Look Once to Hear: Target Speech Hearing with Noisy ExamplesBandhav Veluri, Malek Itani, Tuochao Chen, Takuya Yoshioka 等CHI 2024 · 被引用 25 次
- Spatial Speech Translation: Translating Across Space With Binaural HearablesTuochao Chen, Qirui Wang, Runlin He, Shyamnath GollakotaCHI 2025 · 被引用 5 次
- Wireless Hearables With Programmable Speech AI AcceleratorsMalek Itani, Tuochao Chen, Arun Raghavan, Gavriel Kohlberg 等MobiCom 2025 · 被引用 4 次
- A Survey of Earable Technology: Trends, Tools, and the Road AheadChangshuo Hu, Qiang Yang, Yang Liu, Tobias Röddiger 等UbiComp 2026 · 被引用 4 次
- VueBuds: Visual Intelligence with Wireless EarbudsMaruchi Kim, Rasya Fawwaz, Zhi Yang Lim, Brinda Moudgalya 等CHI 2026 · 被引用 1 次
它引用的顶会 Paper10
- Co-Separating Sounds of Visual ObjectsRuohan Gao, Kristen GraumanICCV 2019 · 被引用 224 次
- Recursive Visual Sound Separation Using Minus-Plus NetXudong Xu, Bo Dai, Dahua LinICCV 2019 · 被引用 95 次
- EarBuddy: Enabling On-Face Interaction via Wireless EarbudsXuhai Xu, Haitian Shi, Xin Yi, Wenjia Liu 等CHI 2020 · 被引用 91 次
- Rescribe: Authoring and Automatically Editing Audio DescriptionsAmy Pavel, Gabriel Reyes, Jeffrey P. BighamUIST 2020 · 被引用 72 次
- EarSense: earphones as a teeth activity sensorJay Prakash, Zhijian Yang, Yu-Lin Wei, Haitham Hassanieh 等MobiCom 2020 · 被引用 67 次
相关 Paper
- UbiHearo: Bringing Scenario-Aware Sound Guidance to DHH Users with Mobile AgentsFengmin Wu, Sicong Liu, Zimu Zhou, Chenren Xu 等UbiComp 2026
- Dompteur: Taming Audio Adversarial ExamplesThorsten Eisenhofer, Lea Schönherr, Joel Frank, Lars Speckemeier 等USENIX Security 2021 · 被引用 29 次
- Proactive Hearing Assistants that Isolate Egocentric ConversationsGuilin Hu, Malek Itani, Tuochao Chen, Shyamnath GollakotaEMNLP 2025 · 被引用 1 次
- Knowing When to Quit: Probabilistic Early Exits for Speech Separation NetworksKenny Falkær Olsen, Mads Østergaard, Karl Ulbæk, Søren Føns Nielsen 等ICLR 2026 · 被引用 1 次
- Semantic Audio-Visual NavigationChangan Chen, Ziad Al-Halah, Kristen GraumanCVPR 2021
