MetricGAN-OKD: Multi-Metric Optimization of MetricGAN via Online Knowledge Distillation for Speech Enhancement
Wooseok Shin, Byung Hoon Lee, Jin Sob Kim, Hyun Joon Park, Sung Won Han
Abstract
In speech enhancement, MetricGAN-based approaches reduce the discrepancy between the L p loss and evaluation metrics by utilizing a non-differentiable evaluation metric as the objective function. However, optimizing multiple metrics simultaneously remains challenging owing to the problem of confusing gradient directions. In this paper, we propose an effective multi-metric optimization method in MetricGAN via online knowledge distillation-MetricGAN-OKD. MetricGAN-OKD, which consists of multiple generators and target metrics, related by a one-to-one correspondence, enables generators to learn with respect to a single metric reliably while improving performance with respect to other metrics by mimicking other generators. Experimental results on speech enhancement and listening enhancement tasks reveal that the proposed method significantly improves performance in terms of multiple metrics compared to existing multi-metric optimization methods. Further, the good performance of MetricGAN-OKD is explained in terms of network generalizability and correlation between metrics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on4
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- When Vision Transformers Outperform ResNets without Pre-training or Strong Data AugmentationsXiangning Chen, Cho-Jui Hsieh, Boqing GongICLR 2022 · 388 citations
- PHASEN: A Phase-and-Harmonics-Aware Speech Enhancement NetworkDacheng Yin, Chong Luo, Zhiwei Xiong, Wenjun ZengAAAI 2020 · 387 citations
- Online Knowledge Distillation via Collaborative LearningQiushan Guo, Xinjiang Wang, Yichao Wu, Zhipeng Yu et al.CVPR 2020
Related papers
- From Continuous to Discrete: Cross-Domain Collaborative General Speech Enhancement via Hierarchical Language ModelsZhaoxi Mu, Rilin Chen, Andong Li, Meng Yu et al.ACM MM 2025
- A Spectral Energy Distance for Parallel Speech SynthesisAlexey A. Gritsenko, Tim Salimans, Rianne van den Berg, Jasper Snoek et al.NeurIPS 2020 · 89 citations
- DMOSpeech: Direct Metric Optimization via Distilled Diffusion Model in Zero-Shot Speech SynthesisYinghao Aaron Li, Rithesh Kumar, Zeyu JinICML 2025
- Agree to Disagree: Adaptive Ensemble Knowledge Distillation in Gradient SpaceShangchen Du, Shan You, Xiaojie Li, Jianlong Wu et al.NeurIPS 2020 · 144 citations
- Multi-SpectroGAN: High-Diversity and High-Fidelity Spectrogram Generation with Adversarial Style Combination for Speech SynthesisSang-Hoon Lee, Hyun-Wook Yoon, Hyeong-Rae Noh, Ji-Hoon Kim et al.AAAI 2021 · 60 citations
