SafeEar: Content Privacy-Preserving Audio Deepfake Detection
Xinfeng Li, Kai Li, Yifan Zheng, Chen Yan, Xiaoyu Ji, Wenyuan Xu
Abstract
Text-to-Speech (TTS) and Voice Conversion (VC) models have exhibited remarkable performance in generating realistic and natural audio. However, their dark side, audio deepfake poses a significant threat to both society and individuals. Existing countermeasures largely focus on determining the genuineness of speech based on complete original audio recordings, which however often contain private content. This oversight may refrain deepfake detection from many applications, particularly in scenarios involving sensitive information like business secrets. In this paper, we propose SafeEar, a novel framework that aims to detect deepfake audios without relying on accessing the speech content within. Our key idea is to devise a neural audio codec into a novel decoupling model that well separates the semantic and acoustic information from audio samples, and only use the acoustic information (e.g., prosody and timbre) for deepfake detection. In this way, no semantic content will be exposed to the detector. To overcome the challenge of identifying diverse deepfake audio without semantic clues, we enhance our deepfake detector with real-world codec augmentation. Extensive experiments conducted on four benchmark datasets demonstrate SafeEar's effectiveness in detecting various deepfake techniques with an equal error rate (EER) down to 2.02%. Simultaneously, it shields five-language speech content from being deciphered by both machine and human auditory analysis, demonstrated by word error rates (WERs) all above 93.93% and our user study. Furthermore, our benchmark constructed for anti-deepfake and anti-content recovery evaluation helps provide a basis for future research in the realms of audio privacy preservation and deepfake detection.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 72cecb7d-9edc-4727-959b-8d316b5b66a2Cited by top-tier papers2
- AudioTrust: Benchmarking The Multifaceted Trustworthiness of Audio Large Language ModelsKai Li, Can Shen, Yile Liu, Jirui Han et al.ICLR 2026 · 17 citations
- Omni-Fake: Benchmarking Unified Multimodal Social Media Deepfake DetectionTianxiao Li, Zhenglin Huang, Haiquan Wen, Yiwei He et al.CVPR 2026 · 5 citations
Builds on13
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 2,890 citations
- Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-SpeechJaehyeon Kim, Jungil Kong, Juhee SonICML 2021 · 1,267 citations
- FastSpeech 2: Fast and High-Quality End-to-End Text to SpeechYi Ren, Chenxu Hu, Xu Tan, Tao Qin et al.ICLR 2021 · 513 citations
- Speech Separation Using an Asynchronous Fully Recurrent Convolutional Neural NetworkXiaolin Hu, Kai Li, Weiyi Zhang, Yi Luo et al.NeurIPS 2021 · 74 citations
- The Catcher in the Field: A Fieldprint based Spoofing Detection for Text-Independent Speaker VerificationChen Yan, Yan Long, Xiaoyu Ji, Wenyuan XuCCS 2019 · 62 citations
Related papers
- Joint Audio-Visual Deepfake DetectionYipin Zhou, Ser-Nam LimICCV 2021 · 232 citations
- SepVAMark: Deep Separable Visual-Audio Fusion Watermarking for Source Tracing and Deepfake DetectionChuan Zhang, Zihan Li, Zihao Xu, Xuhao Ren et al.ACM MM 2025 · 2 citations
- Audio Deepfake Detection with Self-Supervised XLS-R and SLS ClassifierQishan Zhang, Shuangbing Wen, Tao HuACM MM 2024 · 54 citations
- SafeSpeech: Robust and Universal Voice Protection Against Malicious Speech SynthesisZhisheng Zhang, Derui Wang, Qianyi Yang, Pengyang Huang et al.USENIX Security 2025
- WhiADD: Semantic-Acoustic Fusion for Robust Audio Deepfake DetectionJianqiao Cui, Bingyao Yu, Qihao Wang, Fei Meng et al.ACM MM 2025 · 1 citation
