Enkidu: Universal Frequential Perturbation for Real-Time Audio Privacy Protection against Voice Deepfakes
Zhou Feng, Jiahao Chen, Chunyi Zhou, Yuwen Pu, Qingming Li, Tianyu Du, Shouling Ji
Abstract
The rise of advanced voice deepfake technologies has raised serious concerns over user audio privacy, as malicious actors increasingly exploit publicly available voice data to generate convincing fake audio for malicious purposes such as identity theft, financial fraud and misinformation campaigns. While existing defense methods offer partial protection, they suffer from critical limitations, including weak adaptability to unseen user data, poor scalability to long audio, regid reliance on white-box knowledge and high computational and temporal costs to encryption process. Therefore, to defend against personalized voice deepfake threats, we propose Enkidu, a novel user-oriented privacy-preserving framework that leverages universal frequential perturbations generated through black-box knowledge and few-shot training on a small amount of user samples. These high-malleablity frequency-domain noise patches enable real-time, lightweight protection with strong generalization across variable-length audio and robust resistance against voice deepfake attacks-all while preserving high perceptual and intelligible audio quality. Notably, Enkidu achieves over 50-200× processing memory efficiency (requiring only 0.004 GB) and over 3-7000× runtime efficiency (real-time coefficient as low as 0.004) compared to six SOTA countermeasures. Extensive experiments across six mainstream Text-to-Speech (TTS) models and five cutting-edge Automated Speaker Verification (ASV) models demonstrate the effectiveness, transferability, and practicality of Enkidu in defending against voice deepfakes and adaptive attacks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on10
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- Glow-TTS: A Generative Flow for Text-to-Speech via Monotonic Alignment SearchJaehyeon Kim, Sungwon Kim, Jungil Kong, Sungroh YoonNeurIPS 2020 · 663 citations
- YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot Voice Conversion for EveryoneEdresson Casanova, Julian Weber, Christopher Dane Shulby, Arnaldo Cândido Júnior et al.ICML 2022 · 602 citations
- FastSpeech 2: Fast and High-Quality End-to-End Text to SpeechYi Ren, Chenxu Hu, Xu Tan, Tao Qin et al.ICLR 2021 · 513 citations
Related papers
- PhonoFence: A Cross-Task Defense Framework for DeepFake via Phoneme-Level Adversarial PerturbationsZhaolin Wei, Xiuwen Shi, Dengpan Ye, Yuhan Lin et al.ACM MM 2025
- SafeSpeech: Robust and Universal Voice Protection Against Malicious Speech SynthesisZhisheng Zhang, Derui Wang, Qianyi Yang, Pengyang Huang et al.USENIX Security 2025
- VoiceBlock: Privacy through Real-Time Adversarial Attacks with Audio-to-Audio ModelsPatrick O'Reilly, Andreas Bugler, Keshav Bhandari, Max Morrison et al.NeurIPS 2022 · 18 citations
- De-AntiFake: Rethinking the Protective Perturbations Against Voice Cloning AttacksWei Fan, Kejiang Chen, Chang Liu, Weiming Zhang et al.ICML 2025
- V-Cloak: Intelligibility-, Naturalness- & Timbre-Preserving Real-Time Voice AnonymizationJiangyi Deng, Fei Teng, Yanjiao Chen, Xiaofu Chen et al.USENIX Security 2023
