Multi-level SSL Feature Gating for Audio Deepfake Detection
Hoan My Tran, Damien Lolive, Aghilas Sini, Arnaud Delhay, Pierre-François Marteau, David Guennec
摘要
Recent advancements in generative AI, particularly in speech synthesis, have enabled the generation of highly natural-sounding synthetic speech that closely mimics human voices. While these innovations hold promise for applications like assistive technologies, they also pose significant risks, including misuse for fraudulent activities, identity theft, and security threats. Current research on spoofing detection countermeasures remains limited by generalization to unseen deepfake attacks and languages. To address this, we propose a gating mechanism extracting relevant feature from the speech foundation XLS-R model as a front-end feature extractor. For downstream back-end classifier, we employ Multi-kernel gated Convolution (MultiConv) to capture both local and global speech artifacts. Additionally, we introduce Centered Kernel Alignment (CKA) as a similarity metric to enforce diversity in learned features across different MultiConv layers. By integrating CKA with our gating mechanism, we hypothesize that each component helps improving the learning of distinct synthetic speech patterns. Experimental results demonstrate that our approach achieves state-of-the-art performance on in-domain benchmarks while generalizing robustly to out-of-domain datasets, including multilingual speech samples. This underscores its potential as a versatile solution for detecting evolving speech deepfake threats.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper7
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- Pay Attention to MLPsHanxiao Liu, Zihang Dai, David R. So, Quoc V. LeNeurIPS 2021 · 被引用 912 次
- Audio Deepfake Detection with Self-Supervised XLS-R and SLS ClassifierQishan Zhang, Shuangbing Wen, Tao HuACM MM 2024 · 被引用 54 次
- Transferring Audio Deepfake Detection Capability across LanguagesZhongjie Ba, Qing Wen, Peng Cheng, Yuwei Wang 等WWW 2023 · 被引用 34 次
- MCPNet: An Interpretable Classifier via Multi-Level Concept PrototypesBor-Shiun Wang, Chien-Yi Wang, Wei-Chen ChiuCVPR 2024 · 被引用 11 次
相关 Paper
- SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation MethodsWen Huang, Yanmei Gu, Zhiming Wang, Huijia Zhu 等ACL 2025
- AntiFake: Using Adversarial Audio to Prevent Unauthorized Speech SynthesisZhiyuan Yu, Shixuan Zhai, Ning ZhangCCS 2023 · 被引用 29 次
- Generalizable Audio Deepfake Detection via Risk-Aware Style Alignment and Structural Empirical Risk MinimizationMingru Yang, Yanmei Gu, Qianhua He, Peirong Zhang 等ACM MM 2025 · 被引用 1 次
- Improving Generalization for AI-Synthesized Voice DetectionHainan Ren, Li Lin, Chun-Hao Liu, Xin Wang 等AAAI 2025 · 被引用 13 次
- Phoneme-Level Feature Discrepancies: A Key to Detecting Sophisticated Speech DeepfakesKuiyuan Zhang, Zhongyun Hua, Rushi Lan, Yushu Zhang 等AAAI 2025 · 被引用 5 次
