PGSS: Pitch-Guided Speech Separation
Xiang Li, Yiwen Wang, Yifan Sun, Xihong Wu, Jing Chen
Abstract
Monaural speech separation aims to separate concurrent speakers from a single-microphone mixture recording. Inspired by the effect of pitch priming in auditory scene analysis (ASA) mechanisms, a novel pitch-guided speech separation framework is proposed in this work. The prominent advantage of this framework is that both the permutation problem and the unknown speaker number problem existing in general models can be avoided by using pitch contours as the primary means to guide the target speaker. In addition, adversarial training is applied, instead of a traditional time-frequency mask, to improve the perceptual quality of separated speech. Specifically, the proposed framework can be divided into two phases: pitch extraction and speech separation. The former aims to extract pitch contour candidates for each speaker from the mixture, modeling the bottom-up process in ASA mechanisms. Any pitch contour can be selected as the condition in the second phase to separate the corresponding speaker, where a conditional generative adversarial network (CGAN) is applied. The second phase models the effect of pitch priming in ASA. Experiments on the WSJ0-2mix corpus reveal that the proposed approaches can achieve higher pitch extraction accuracy and better separation performance, compared to the baseline models, and have the potential to be applied to SOTA architectures.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d51f483c-235d-44a3-a04b-936f2ccf6ee7Cited by top-tier papers2
- OmniSep: Unified Omni-Modality Sound Separation with Query-MixupXize Cheng, Siqi Zheng, Zehan Wang, Minghui Fang et al.ICLR 2025
- AlignSep: Temporally-Aligned Video-Queried Sound Separation with Flow MatchingXize Cheng, Chenyuhao Wen, Slytherin Wang, Yongqi Wang et al.ICLR 2026
Related papers
- Multi-SpectroGAN: High-Diversity and High-Fidelity Spectrogram Generation with Adversarial Style Combination for Speech SynthesisSang-Hoon Lee, Hyun-Wook Yoon, Hyeong-Rae Noh, Ji-Hoon Kim et al.AAAI 2021 · 60 citations
- Unsupervised Single-Channel Audio Separation with Diffusion Source PriorsRunwu Shi, Chang Li, Jiang Wang, Rui Zhang et al.AAAI 2026
- UNSSOR: Unsupervised Neural Speech Separation by Leveraging Over-determined Training MixturesZhong-Qiu Wang, Shinji WatanabeNeurIPS 2023 · 24 citations
- Synthesising Audio Adversarial Examples for Automatic Speech RecognitionXinghua Qu, Pengfei Wei, Mingyong Gao, Zhu Sun et al.KDD 2022 · 7 citations
- The Cone of Silence: Speech Separation by LocalizationTeerapat Jenrungrot, Vivek Jayaram, Steven M. Seitz, Ira Kemelmacher-ShlizermanNeurIPS 2020 · 70 citations
