The Cone of Silence: Speech Separation by Localization
Teerapat Jenrungrot, Vivek Jayaram, Steven M. Seitz, Ira Kemelmacher-Shlizerman
摘要
Given a multi-microphone recording of an unknown number of speakers talking concurrently, we simultaneously localize the sources and separate the individual speakers. At the core of our method is a deep network, in the waveform domain, which isolates sources within an angular region , given an angle of interest and angular window size . By exponentially decreasing , we can perform a binary search to localize and separate all sources in logarithmic time. Our algorithm allows for an arbitrary number of potentially moving speakers at test time, including more speakers than seen during training. Experiments demonstrate state-of-the-art performance for both source separation and source localization, particularly in high levels of background noise.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- MESH2IR: Neural Acoustic Impulse Response Generator for Complex 3D ScenesAnton Ratnarajah, Zhenyu Tang, Rohith Aralikatti, Dinesh ManochaACM MM 2022 · 被引用 35 次
- Look Once to Hear: Target Speech Hearing with Noisy ExamplesBandhav Veluri, Malek Itani, Tuochao Chen, Takuya Yoshioka 等CHI 2024 · 被引用 25 次
- UNSSOR: Unsupervised Neural Speech Separation by Leveraging Over-determined Training MixturesZhong-Qiu Wang, Shinji WatanabeNeurIPS 2023 · 被引用 24 次
- GWA: A Large High-Quality Acoustic Dataset for Audio ProcessingZhenyu Tang, Rohith Aralikatti, Anton Jeran Ratnarajah, Dinesh ManochaSIGGRAPH 2022 · 被引用 23 次
- Hybrid Neural Networks for On-Device Directional HearingAnran Wang, Maruchi Kim, Hao Zhang, Shyamnath GollakotaAAAI 2022 · 被引用 18 次
它引用的顶会 Paper2
相关 Paper
- Egocentric Deep Multi-Channel Audio-Visual Active Speaker LocalizationHao Jiang, Calvin Murdock, Vamsi Krishna IthapuCVPR 2022 · 被引用 39 次
- mmWave-Aided Unified Speech Enhancement and Separation without Speaker Count PriorDachao Han, Teng Huang, Han Ding, Cui Zhao 等INFOCOM 2026 · 被引用 1 次
- Filter-Recovery Network for Multi-Speaker Audio-Visual Speech SeparationHaoyue Cheng, Zhaoyang Liu, Wayne Wu, Limin WangICLR 2023
- Learning to Separate Voices by Spatial RegionsAlan Xu, Romit Roy ChoudhuryICML 2022 · 被引用 18 次
- DeepRange: Acoustic Ranging via Deep LearningWenguang Mao, Wei Sun, Mei Wang, Lili QiuUbiComp 2021 · 被引用 49 次
