Audio Deepfake Detection with Self-Supervised XLS-R and SLS Classifier
Qishan Zhang, Shuangbing Wen, Tao Hu
Abstract
Generative AI technologies, including text-to-speech (TTS) and voice conversion (VC), frequently become indistinguishable from genuine samples, posing challenges for individuals in discerning between real and synthetic content. This indistinguishability undermines trust in media, and the arbitrary cloning of personal voice signals presents significant challenges to privacy and security. In the field of deepfake audio detection, the majority of models achieving higher detection accuracy currently employ self-supervised pre-trained models. However, with the ongoing development of deepfake audio generation algorithms, maintaining high discrimination accuracy against new algorithms grows more challenging. To enhance the sensitivity of deepfake audio features, we propose a deepfake audio detection model that incorporates an SLS (Sensitive Layer Selection) module. Specifically, utilizing the pre-trained XLS-R enables our model to extract diverse audio features from its various layers, each providing distinct discriminative information. Utilizing the SLS classifier, our model captures sensitive contextual information across different layer levels of audio features, effectively employing this information for fake audio detection. Experimental results show that our method achieves state-of-the-art (SOTA) performance on both the ASVspoof 2021 DF and In-the-Wild datasets, with a specific Equal Error Rate (EER) of 1.92% on the ASVspoof 2021 DF dataset and 7.46% on the In-the-Wild dataset. Codes and data can be found at https://github.com/QiShanZhang/SLSforADD.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get eff204bb-4f78-466e-8b9d-a54489a2f4feCited by top-tier papers6
- ALLM4ADD: Unlocking the Capabilities of Audio Large Language Models for Audio Deepfake DetectionHao Gu, Jiangyan Yi, Chenglong Wang, Jianhua Tao et al.ACM MM 2025 · 5 citations
- Multi-level SSL Feature Gating for Audio Deepfake DetectionHoan My Tran, Damien Lolive, Aghilas Sini, Arnaud Delhay et al.ACM MM 2025 · 3 citations
- HyperPotter: Spell the Charm of High-Order Interactions in Audio Deepfake DetectionQing Wen, Haohao Li, Zhongjie Ba, Peng Cheng et al.ICML 2026 · 1 citation
- SONAR: Spectral‑Contrastive Audio Residuals for Generalizable Deepfake DetectionIdo Nitzan Hidekel, Gal Lifshitz, Khen Cohen, Dan RavivICML 2026 · 1 citation
- SiFMimicEvader: Evading Fake Voice Detection with Adversarial Neural Mimicry AttacksXuan Hai, Xin Liu, Zihao Zhang, Ziyao Yu et al.ACM MM 2025
Related papers
- Phoneme-Level Feature Discrepancies: A Key to Detecting Sophisticated Speech DeepfakesKuiyuan Zhang, Zhongyun Hua, Rushi Lan, Yushu Zhang et al.AAAI 2025 · 5 citations
- SafeEar: Content Privacy-Preserving Audio Deepfake DetectionXinfeng Li, Kai Li, Yifan Zheng, Chen Yan et al.CCS 2024 · 26 citations
- Detect All-Type Deepfake Audio: Wavelet Prompt Tuning for Enhanced Auditory PerceptionYuankun Xie, Ruibo Fu, Xiaopeng Wang, Zhiyong Wang et al.AAAI 2026 · 9 citations
- SLIM: Style-Linguistics Mismatch Model for Generalized Audio Deepfake DetectionYi Zhu, Surya Koppisetti, Trang Tran, Gaurav BharajNeurIPS 2024 · 41 citations
- Investigating Self-Supervised Representations for Audio-Visual Deepfake DetectionDragos-Alexandru Boldisor, Stefan Smeu, Dan Oneata, Elisabeta OneataCVPR 2026 · 2 citations
