Multi-View Collaborative Learning Network for Speech Deepfake Detection
Kuiyuan Zhang, Zhongyun Hua, Rushi Lan, Yifang Guo, Yushu Zhang, Guoai Xu
摘要
As deep learning techniques advance rapidly, deepfake speech synthesized through text-to-speech or voice conversion networks is becoming increasingly realistic, posing significant challenges for detection and raising potential threats to social security. This growing realism has prompted extensive research in speech deepfake detection. However, current detection methods primarily focus on extracting features from either the raw waveform or the spectrogram, often overlooking the valuable correspondences between these two modalities that could enhance the detection of previously unseen types of deepfakes. In this work, we propose a multi-view collaborative learning network for speech deepfake detection, which jointly learns robust speech representations from both raw waveforms and spectrograms. Specifically, we first design a Dual-Branch Contrastive Learning (DBCL) framework for learning different view features. DBCL consists of two branches that learn representations from the raw waveform or the spectrogram and utilizes contrastive learning to enhance inter- and inner-view correlations. Additionally, we introduce a Waveform-Spectrogram Fusion Module (WSFM) to exchange multi-view information for collaborative learning. In the feature learning process, WSFM converts features between views and merges them adaptively using waveform-spectrogram cross-attention. The final detection is conducted based on the concatenation of the waveform and spectrogram features. We conduct extensive experiments on four benchmark deepfake speech detection datasets, and the experimental results demonstrate that our method can achieve better detection performance than current state-of-the-art detection methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-SpeechJaehyeon Kim, Jungil Kong, Juhee SonICML 2021 · 被引用 1,267 次
- Grad-TTS: A Diffusion Probabilistic Model for Text-to-SpeechVadim Popov, Ivan Vovk, Vladimir Gogoryan, Tasnima Sadekova 等ICML 2021 · 被引用 715 次
- YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot Voice Conversion for EveryoneEdresson Casanova, Julian Weber, Christopher Dane Shulby, Arnaldo Cândido Júnior 等ICML 2022 · 被引用 602 次
- Guided-TTS: A Diffusion Model for Text-to-Speech via Classifier GuidanceHeeseung Kim, Sungwon Kim, Sungroh YoonICML 2022 · 被引用 133 次
- HierSpeech: Bridging the Gap between Text and Speech by Hierarchical Variational Inference using Self-supervised Representations for Speech SynthesisSang-Hoon Lee, Seung-Bin Kim, Ji-Hyun Lee, Eunwoo Song 等NeurIPS 2022 · 被引用 81 次
相关 Paper
- AVFF: Audio-Visual Feature Fusion for Video Deepfake DetectionTrevine Oorloff, Surya Koppisetti, Nicolò Bonettini, Divyaraj Solanki 等CVPR 2024 · 被引用 51 次
- Multi-modal Deepfake Detection via Multi-task Audio-Visual Prompt LearningHui Miao, Yuanfang Guo, Zeming Liu, Yunhong WangAAAI 2025 · 被引用 8 次
- Joint Audio-Visual Deepfake DetectionYipin Zhou, Ser-Nam LimICCV 2021 · 被引用 232 次
- Rethinking Fake Speech Detection: A Generalized Framework Leveraging Spectrogram MagnitudeZihao Liu, Aobo Chen, Yan Zhang, Wensheng Zhang 等NDSS 2026
- SpecXNet: A Dual-Domain Convolutional Network for Robust Deepfake DetectionInzamamul Alam, Md Tanvir Islam, Simon S. WooACM MM 2025 · 被引用 6 次
