Learning to Separate Voices by Spatial Regions
Alan Xu, Romit Roy Choudhury
摘要
We consider the problem of audio voice separation for binaural applications, such as earphones and hearing aids. While today's neural networks perform remarkably well (separating sources with 2 microphones) they assume a known or fixed maximum number of sources, K. Moreover, today's models are trained in a supervised manner, using training data synthesized from generic sources, environments, and human head shapes. This paper intends to relax both these constraints at the expense of a slight alteration in the problem definition. We observe that, when a received mixture contains too many sources, it is still helpful to separate them by region, i.e., isolating signal mixtures from each conical sector around the user's head. This requires learning the fine-grained spatial properties of each region, including the signal distortions imposed by a person's head. We propose a two-stage self-supervised framework in which overheard voices from earphones are pre-processed to extract relatively clean personalized signals, which are then used to train a region-wise separation model. Results show promising performance, underscoring the importance of personalization over a generic supervised approach. (audio samples available at our project website: https://uiuc-earable-computing.github.io/binaural/. We believe this result could help real-world applications in selective hearing, noise cancellation, and audio augmented reality.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Spatial Speech Translation: Translating Across Space With Binaural HearablesTuochao Chen, Qirui Wang, Runlin He, Shyamnath GollakotaCHI 2025 · 被引用 5 次
- ArrayDPS: Unsupervised Blind Speech Separation with a Diffusion PriorZhongweiyang Xu, Xulin Fan, Zhong-Qiu Wang, Xilin Jiang 等ICML 2025
- For Human Ears Only: Preventing Automated Monitoring on Voice DataIrtaza Shahid, Nirupam RoyUSENIX Security 2025
它引用的顶会 Paper4
- Unsupervised Sound Separation Using Mixture Invariant TrainingScott Wisdom, Efthymios Tzinis, Hakan Erdogan, Ron J. Weiss 等NeurIPS 2020 · 被引用 227 次
- Voice Separation with an Unknown Number of Multiple SpeakersEliya Nachmani, Yossi Adi, Lior WolfICML 2020 · 被引用 186 次
- The Cone of Silence: Speech Separation by LocalizationTeerapat Jenrungrot, Vivek Jayaram, Steven M. Seitz, Ira Kemelmacher-ShlizermanNeurIPS 2020 · 被引用 70 次
- Personalizing head related transfer functions for earablesZhijian Yang, Romit Roy ChoudhurySIGCOMM 2021 · 被引用 26 次
相关 Paper
- DeepEar: Sound Localization with Binaural MicrophonesQiang Yang, Yuanqing ZhengINFOCOM 2022 · 被引用 8 次
- Neural Synthesis of Binaural Speech From Mono AudioAlexander Richard, Dejan Markovic, Israel D. Gebru, Steven Krenn 等ICLR 2021 · 被引用 73 次
- Proactive Hearing Assistants that Isolate Egocentric ConversationsGuilin Hu, Malek Itani, Tuochao Chen, Shyamnath GollakotaEMNLP 2025 · 被引用 1 次
- Learning Spatial Features from Audio-Visual Correspondence in Egocentric VideosSagnik Majumder, Ziad Al-Halah, Kristen GraumanCVPR 2024 · 被引用 3 次
- Localize to Binauralize: Audio Spatialization from Visual Sound Source LocalizationKranthi Kumar Rachavarapu, Aakanksha, Vignesh Sundaresha, A. N. RajagopalanICCV 2021 · 被引用 27 次
