PhonoFence: A Cross-Task Defense Framework for DeepFake via Phoneme-Level Adversarial Perturbations
Zhaolin Wei, Xiuwen Shi, Dengpan Ye, Yuhan Lin, Zhigang Wang, Jiacheng Deng, Ziyi Liu, Long Tang
Abstract
Recent advances in deep neural networks and generative AI have enabled the creation of highly realistic synthetic speech, raising significant concerns about the misuse of deepfake in deception and fraud. This paper addresses the generalization challenge of active defense against cross-task audio deepfake under black-box conditions. We systematically analyze the common characteristics of state-of-the-art acoustic synthesis models and propose PhonoFence, an active defense framework that introduces fine-grained phoneme-level adversarial perturbations to prevent unauthorized synthesis. PhonoFence employs a dual-domain strategy to perturb both the time domain and frequency domain via an iterative cross-training framework, leveraging complementary acoustic features to enhance the generalization of perturbations. To further improve the transferability of perturbations, we ensemble speaker encoders with a novel Multi-Middle-Layer loss. Additionally, a psychoacoustic masking algorithm is employed to enhance the perceptual quality of protected speech and conceal perturbations. Extensive experiments on leading acoustic synthesis models demonstrate that PhonoFence reduces identity similarity and word error rate to 17.91% and 49.77%, respectively, achieving relative improvements of 7.97% and 9.67% over the best existing methods. To assess the effectiveness of our method in real-world, we test PhonoFence in commercial speaker recognition system, where it reduces deepfake attack success rates by 73.49%. Moreover, PhonoFence shows strong robustness against adaptive attacks involving compression, denoising, and re-recording.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 6c498353-34b3-4609-b47c-57c8d1055e19Cited by top-tier papers1
Ask how each one uses itRelated papers
- AntiFake: Using Adversarial Audio to Prevent Unauthorized Speech SynthesisZhiyuan Yu, Shixuan Zhai, Ning ZhangCCS 2023 · 29 citations
- SafeSpeech: Robust and Universal Voice Protection Against Malicious Speech SynthesisZhisheng Zhang, Derui Wang, Qianyi Yang, Pengyang Huang et al.USENIX Security 2025
- Enkidu: Universal Frequential Perturbation for Real-Time Audio Privacy Protection against Voice DeepfakesZhou Feng, Jiahao Chen, Chunyi Zhou, Yuwen Pu et al.ACM MM 2025 · 1 citation
- PhoneyTalker: An Out-of-the-Box Toolkit for Adversarial Example Attack on Speaker RecognitionMeng Chen, Li Lu, Zhongjie Ba, Kui RenINFOCOM 2022 · 13 citations
- De-AntiFake: Rethinking the Protective Perturbations Against Voice Cloning AttacksWei Fan, Kejiang Chen, Chang Liu, Weiming Zhang et al.ICML 2025
