Detecting Voice Cloning Attacks via Timbre Watermarking
Chang Liu, Jie Zhang, Tianwei Zhang, Xi Yang, Weiming Zhang, Nenghai Yu
Abstract
Nowadays, it is common to release audio content to the public. However, with the rise of voice cloning technology, attackers have the potential to easily impersonate a specific person by utilizing his publicly released audio without any permission. Therefore, it becomes significant to detect any potential misuse of the released audio content and protect its timbre from being impersonated. To this end, we introduce a novel concept,"Timbre Watermarking", which embeds watermark information into the target individual's speech, eventually defeating the voice cloning attacks. To ensure the watermark is robust to the voice cloning model's learning process, we design an end-to-end voice cloning-resistant detection framework. The core idea of our solution is to embed and extract the watermark in the frequency domain in a temporally invariant manner. To acquire generalization across different voice cloning attacks, we modulate their shared process and integrate it into our framework as a distortion layer. Experiments demonstrate that the proposed timbre watermarking can defend against different voice cloning attacks, exhibit strong resistance against various adaptive attacks (e.g., reconstruction-based removal attacks, watermark overwriting attacks), and achieve practicality in real-world services such as PaddleSpeech, Voice-Cloning-App, and so-vits-svc. In addition, ablation studies are also conducted to verify the effectiveness of our design. Some audio samples are available at https://timbrewatermarking.github.io/samples.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers15
- GS-Hider: Hiding Messages into 3D Gaussian SplattingXuanyu Zhang, Jiarui Meng, Runyi Li, Zhipei Xu et al.NeurIPS 2024 · 43 citations
- GROOT: Generating Robust Watermark for Diffusion-Model-Based Audio SynthesisWeizhi Liu, Yue Li, Dongdong Lin, Hui Tian et al.ACM MM 2024 · 12 citations
- Speech Watermarking with Discrete Intermediate RepresentationsShengpeng Ji, Ziyue Jiang, Jialong Zuo, Minghui Fang et al.AAAI 2025 · 10 citations
- Yours or Mine? Overwriting Attacks Against Neural Audio WatermarkingLingfeng Yao, Chenpei Huang, Shengyao Wang, Junpei Xue et al.AAAI 2026 · 5 citations
- Robust Distortion-Free Watermark for Autoregressive Audio Generation ModelsYihan Wu, Georgios Milis, Ruibo Chen, Heng HuangNeurIPS 2025 · 5 citations
Builds on11
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 2,890 citations
- Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-SpeechJaehyeon Kim, Jungil Kong, Juhee SonICML 2021 · 1,267 citations
- AudioLDM: Text-to-Audio Generation with Latent Diffusion ModelsHaohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei et al.ICML 2023 · 773 citations
- FastSpeech 2: Fast and High-Quality End-to-End Text to SpeechYi Ren, Chenxu Hu, Xu Tan, Tao Qin et al.ICLR 2021 · 513 citations
- CommanderSong: A Systematic Approach for Practical Adversarial Voice RecognitionXuejing Yuan, Yuxuan Chen, Yue Zhao, Yunhui Long et al.USENIX Security 2018 · 389 citations
Related papers
- E2E-VGuard: Adversarial Prevention for Production LLM-based End-To-End Speech SynthesisZhisheng Zhang, Derui Wang, Yifan Mi, Zhiyong Wu et al.NeurIPS 2025
- AudioMarkNet: Audio Watermarking for Deepfake Speech DetectionWei Zong, Yang-Wai Chow, Willy Susilo, Joonsang Baek et al.USENIX Security 2025
- De-AntiFake: Rethinking the Protective Perturbations Against Voice Cloning AttacksWei Fan, Kejiang Chen, Chang Liu, Weiming Zhang et al.ICML 2025
- DeAR: A Deep-Learning-Based Audio Re-recording Resilient WatermarkingChang Liu, Jie Zhang, Han Fang, Zehua Ma et al.AAAI 2023 · 67 citations
- XAttnMark: Learning Robust Audio Watermarking with Cross-AttentionYixin Liu, Lie Lu, Jihui Jin, Lichao Sun et al.ICML 2025
