V-Cloak: Intelligibility-, Naturalness- & Timbre-Preserving Real-Time Voice Anonymization
Jiangyi Deng, Fei Teng, Yanjiao Chen, Xiaofu Chen, Zhaohui Wang, Wenyuan Xu
摘要
Voice data generated on instant messaging or social media applications contains unique user voiceprints that may be abused by malicious adversaries for identity inference or identity theft. Existing voice anonymization techniques, e.g., signal processing and voice conversion/synthesis, suffer from degradation of perceptual quality. In this paper, we develop a voice anonymization system, named V-CLOAK, which attains real-time voice anonymization while preserving the intelligibility, naturalness and timbre of the audio. Our designed anonymizer features a one-shot generative model that modulates the features of the original audio at different frequency levels. We train the anonymizer with a carefully-designed loss function. Apart from the anonymity loss, we further incorporate the intelligibility loss and the psychoacoustics-based naturalness loss. The anonymizer can realize untargeted and targeted anonymization to achieve the anonymity goals of unidentifiability and unlinkability. We have conducted extensive experiments on four datasets, i.e., LibriSpeech (English), AISHELL (Chinese), Common-Voice (French) and CommonVoice (Italian), five Automatic Speaker Verification (ASV) systems (including two DNNbased, two statistical and one commercial ASV), and eleven Automatic Speech Recognition (ASR) systems (for different languages). Experiment results confirm that V-CLOAK outperforms five baselines in terms of anonymity performance. We also demonstrate that V-CLOAK trained only on the VoxCeleb1 dataset against ECAPA-TDNN ASV and Deep-Speech2 ASR has transferable anonymity against other ASVs and cross-language intelligibility for other ASRs. Furthermore, we verify the robustness of V-CLOAK against various de-noising techniques and adaptive attacks. Hopefully, V-CLOAK may provide a cloak for us in a prism world.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- SafeEar: Content Privacy-Preserving Audio Deepfake DetectionXinfeng Li, Kai Li, Yifan Zheng, Chen Yan 等CCS 2024 · 被引用 26 次
- Certification of Speaker Recognition Models to Additive PerturbationsDmitrii Korzh, Elvir Karimov, Mikhail Pautov, Oleg Y. Rogov 等AAAI 2025 · 被引用 8 次
- MicPro: Microphone-based Voice Privacy ProtectionShilin Xiao, Xiaoyu Ji, Chen Yan, Zhicong Zheng 等CCS 2023 · 被引用 6 次
- Enkidu: Universal Frequential Perturbation for Real-Time Audio Privacy Protection against Voice DeepfakesZhou Feng, Jiahao Chen, Chunyi Zhou, Yuwen Pu 等ACM MM 2025 · 被引用 1 次
- E2E-VGuard: Adversarial Prevention for Production LLM-based End-To-End Speech SynthesisZhisheng Zhang, Derui Wang, Yifan Mi, Zhiyong Wu 等NeurIPS 2025
它引用的顶会 Paper7
- Enhancing Adversarial Example Transferability With an Intermediate Level AttackQian Huang, Isay Katsman, Zeqi Gu, Horace He 等ICCV 2019 · 被引用 293 次
- Who is Real Bob? Adversarial Attacks on Speaker Recognition SystemsGuangke Chen, Sen Chen, Lingling Fan, Xiaoning Du 等S&P 2021 · 被引用 239 次
- AdvPulse: Universal, Synchronization-free, and Targeted Audio Adversarial Attacks via Subsecond PerturbationsZhuohang Li, Yi Wu, Jian Liu, Yingying Chen 等CCS 2020 · 被引用 107 次
- Enabling Fast and Universal Audio Adversarial Attack Using Generative ModelYi Xie, Zhuohang Li, Cong Shi, Jian Liu 等AAAI 2021 · 被引用 77 次
- Black-box Adversarial Attacks on Commercial Speech Platforms with Minimal InformationBaolin Zheng, Peipei Jiang, Qian Wang, Qi Li 等CCS 2021 · 被引用 73 次
相关 Paper
- VoiceCloak: Adversarial Example Enabled Voice De-Identification with Balanced Privacy and UtilityMeng Chen, Li Lu, Junhao Wang, Jiadi Yu 等UbiComp 2023 · 被引用 21 次
- VoiceBlock: Privacy through Real-Time Adversarial Attacks with Audio-to-Audio ModelsPatrick O'Reilly, Andreas Bugler, Keshav Bhandari, Max Morrison 等NeurIPS 2022 · 被引用 18 次
- VoiceCloak: A Multi-Dimensional Defense Framework Against Unauthorized Diffusion-Based Voice CloningQianyue Hu, Junyan Wu, Wei Lu, Xiangyang LuoAAAI 2026
- More Simplicity for Trainers, More Opportunity for Attackers: Black-Box Attacks on Speaker Recognition Systems by Inferring Feature ExtractorYunjie Ge, Pinji Chen, Qian Wang, Lingchen Zhao 等USENIX Security 2024 · 被引用 4 次
- Whispering Under the Eaves: Protecting User Privacy Against Commercial and LLM-powered Automatic Speech Recognition SystemsWeifei Jin, Yuxin Cao, Junjie Su, Derui Wang 等USENIX Security 2025
