Unsupervised Noise Adaptive Speech Enhancement by Discriminator-Constrained Optimal Transport
Hsin-Yi Lin, Huan-Hsin Tseng, Xugang Lu, Yu Tsao
Abstract
This paper presents a novel discriminator-constrained optimal transport network (DOTN) that performs unsupervised domain adaptation for speech enhancement (SE), which is an essential regression task in speech processing. The DOTN aims to estimate clean references of noisy speech in a target domain, by exploiting the knowledge available from the source domain. The domain shift between training and testing data has been reported to be an obstacle to learning problems in diverse fields. Although rich literature exists on unsupervised domain adaptation for classification, the methods proposed, especially in regressions, remain scarce and often depend on additional information regarding the input data. The proposed DOTN approach tactically fuses the optimal transport (OT) theory from mathematical analysis with generative adversarial frameworks, to help evaluate continuous labels in the target domain. The experimental results on two SE tasks demonstrate that by extending the classical OT formulation, our proposed DOTN outperforms previous adversarial domain adaptation frameworks in a purely unsupervised manner.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5e7209f9-9530-4d79-8b30-a79237132201Cited by top-tier papers7
- Large Language Models are Efficient Learners of Noise-Robust Speech RecognitionYuchen Hu, Chen Chen, Chao-Han Huck Yang, Ruizhe Li et al.ICLR 2024 · 41 citations
- Pre-training for Speech Translation: CTC Meets Optimal TransportPhuong-Hang Le, Hongyu Gong, Changhan Wang, Juan Pino et al.ICML 2023 · 33 citations
- Optimal Transport-based Identity Matching for Identity-invariant Facial Expression RecognitionDae Ha Kim, Byung Cheol SongNeurIPS 2022 · 19 citations
- Self-Supervised Visual Acoustic MatchingArjun Somayazulu, Changan Chen, Kristen GraumanNeurIPS 2023 · 17 citations
- Hearing Lips in Noise: Universal Viseme-Phoneme Mapping and Transfer for Robust Audio-Visual Speech RecognitionYuchen Hu, Ruizhe Li, Chen Chen, Chengwei Qin et al.ACL 2023 · 7 citations
Builds on2
Related papers
- Learning Invariant Representation for Unsupervised Image RestorationWenchao Du, Hu Chen, Hongyu YangCVPR 2020
- Deep Domain-Adversarial Image Generation for Domain GeneralisationKaiyang Zhou, Yongxin Yang, Timothy M. Hospedales, Tao XiangAAAI 2020 · 488 citations
- Enhanced Transport Distance for Unsupervised Domain AdaptationMengxue Li, Yiming Zhai, You-Wei Luo, Pengfei Ge et al.CVPR 2020
- Elastic Optimal Transport: Theory, Application, and Empirical EvaluationPei Yang, Yuhang Zhuang, Qi TanICLR 2026
- Unified Optimal Transport Framework for Universal Domain AdaptationWanxing Chang, Ye Shi, Hoang Tuan, Jingya WangNeurIPS 2022 · 118 citations
