Semantic Hearing: Programming Acoustic Scenes with Binaural Hearables
Bandhav Veluri, Malek Itani, Justin Chan, Takuya Yoshioka, Shyamnath Gollakota
Abstract
Imagine being able to listen to the birds chirping in a park without hearing the chatter from other hikers, or being able to block out traffic noise on a busy street while still being able to hear emergency sirens and car honks. We introduce semantic hearing, a novel capability for hearable devices that enables them to, in real-time, focus on, or ignore, specific sounds from real-world environments, while also preserving the spatial cues. To achieve this, we make two technical contributions: 1) we present the first neural network that can achieve binaural target sound extraction in the presence of interfering sounds and background noise, and 2) we design a training methodology that allows our system to generalize to real-world use. Results show that our system can operate with 20 sound classes and that our transformer-based network has a runtime of 6.56 ms on a connected smartphone. In-the-wild evaluation with participants in previously unseen indoor and outdoor scenarios shows that our proof-of-concept system can extract the target sounds and generalize to preserve the spatial cues in its binaural output. Project page with code: https://semantichearing.cs.washington.edu
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d12032ca-600c-41ea-9d70-da32379c694dCited by top-tier papers10
- Look Once to Hear: Target Speech Hearing with Noisy ExamplesBandhav Veluri, Malek Itani, Tuochao Chen, Takuya Yoshioka et al.CHI 2024 · 25 citations
- Spatial Speech Translation: Translating Across Space With Binaural HearablesTuochao Chen, Qirui Wang, Runlin He, Shyamnath GollakotaCHI 2025 · 5 citations
- Wireless Hearables With Programmable Speech AI AcceleratorsMalek Itani, Tuochao Chen, Arun Raghavan, Gavriel Kohlberg et al.MobiCom 2025 · 4 citations
- A Survey of Earable Technology: Trends, Tools, and the Road AheadChangshuo Hu, Qiang Yang, Yang Liu, Tobias Röddiger et al.UbiComp 2026 · 4 citations
- VueBuds: Visual Intelligence with Wireless EarbudsMaruchi Kim, Rasya Fawwaz, Zhi Yang Lim, Brinda Moudgalya et al.CHI 2026 · 1 citation
Builds on10
- Co-Separating Sounds of Visual ObjectsRuohan Gao, Kristen GraumanICCV 2019 · 224 citations
- Recursive Visual Sound Separation Using Minus-Plus NetXudong Xu, Bo Dai, Dahua LinICCV 2019 · 95 citations
- EarBuddy: Enabling On-Face Interaction via Wireless EarbudsXuhai Xu, Haitian Shi, Xin Yi, Wenjia Liu et al.CHI 2020 · 91 citations
- Rescribe: Authoring and Automatically Editing Audio DescriptionsAmy Pavel, Gabriel Reyes, Jeffrey P. BighamUIST 2020 · 72 citations
- EarSense: earphones as a teeth activity sensorJay Prakash, Zhijian Yang, Yu-Lin Wei, Haitham Hassanieh et al.MobiCom 2020 · 67 citations
Related papers
- UbiHearo: Bringing Scenario-Aware Sound Guidance to DHH Users with Mobile AgentsFengmin Wu, Sicong Liu, Zimu Zhou, Chenren Xu et al.UbiComp 2026
- Dompteur: Taming Audio Adversarial ExamplesThorsten Eisenhofer, Lea Schönherr, Joel Frank, Lars Speckemeier et al.USENIX Security 2021 · 29 citations
- Proactive Hearing Assistants that Isolate Egocentric ConversationsGuilin Hu, Malek Itani, Tuochao Chen, Shyamnath GollakotaEMNLP 2025 · 1 citation
- Knowing When to Quit: Probabilistic Early Exits for Speech Separation NetworksKenny Falkær Olsen, Mads Østergaard, Karl Ulbæk, Søren Føns Nielsen et al.ICLR 2026 · 1 citation
- Semantic Audio-Visual NavigationChangan Chen, Ziad Al-Halah, Kristen GraumanCVPR 2021
