Detecting Voice Cloning Attacks via Timbre Watermarking
Chang Liu, Jie Zhang, Tianwei Zhang, Xi Yang, Weiming Zhang, Nenghai Yu
摘要
Nowadays, it is common to release audio content to the public. However, with the rise of voice cloning technology, attackers have the potential to easily impersonate a specific person by utilizing his publicly released audio without any permission. Therefore, it becomes significant to detect any potential misuse of the released audio content and protect its timbre from being impersonated. To this end, we introduce a novel concept,"Timbre Watermarking", which embeds watermark information into the target individual's speech, eventually defeating the voice cloning attacks. To ensure the watermark is robust to the voice cloning model's learning process, we design an end-to-end voice cloning-resistant detection framework. The core idea of our solution is to embed and extract the watermark in the frequency domain in a temporally invariant manner. To acquire generalization across different voice cloning attacks, we modulate their shared process and integrate it into our framework as a distortion layer. Experiments demonstrate that the proposed timbre watermarking can defend against different voice cloning attacks, exhibit strong resistance against various adaptive attacks (e.g., reconstruction-based removal attacks, watermark overwriting attacks), and achieve practicality in real-world services such as PaddleSpeech, Voice-Cloning-App, and so-vits-svc. In addition, ablation studies are also conducted to verify the effectiveness of our design. Some audio samples are available at https://timbrewatermarking.github.io/samples.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- GS-Hider: Hiding Messages into 3D Gaussian SplattingXuanyu Zhang, Jiarui Meng, Runyi Li, Zhipei Xu 等NeurIPS 2024 · 被引用 43 次
- GROOT: Generating Robust Watermark for Diffusion-Model-Based Audio SynthesisWeizhi Liu, Yue Li, Dongdong Lin, Hui Tian 等ACM MM 2024 · 被引用 12 次
- Speech Watermarking with Discrete Intermediate RepresentationsShengpeng Ji, Ziyue Jiang, Jialong Zuo, Minghui Fang 等AAAI 2025 · 被引用 10 次
- Yours or Mine? Overwriting Attacks Against Neural Audio WatermarkingLingfeng Yao, Chenpei Huang, Shengyao Wang, Junpei Xue 等AAAI 2026 · 被引用 5 次
- Robust Distortion-Free Watermark for Autoregressive Audio Generation ModelsYihan Wu, Georgios Milis, Ruibo Chen, Heng HuangNeurIPS 2025 · 被引用 5 次
它引用的顶会 Paper11
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 被引用 2,890 次
- Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-SpeechJaehyeon Kim, Jungil Kong, Juhee SonICML 2021 · 被引用 1,267 次
- AudioLDM: Text-to-Audio Generation with Latent Diffusion ModelsHaohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei 等ICML 2023 · 被引用 773 次
- FastSpeech 2: Fast and High-Quality End-to-End Text to SpeechYi Ren, Chenxu Hu, Xu Tan, Tao Qin 等ICLR 2021 · 被引用 513 次
- CommanderSong: A Systematic Approach for Practical Adversarial Voice RecognitionXuejing Yuan, Yuxuan Chen, Yue Zhao, Yunhui Long 等USENIX Security 2018 · 被引用 389 次
相关 Paper
- E2E-VGuard: Adversarial Prevention for Production LLM-based End-To-End Speech SynthesisZhisheng Zhang, Derui Wang, Yifan Mi, Zhiyong Wu 等NeurIPS 2025
- AudioMarkNet: Audio Watermarking for Deepfake Speech DetectionWei Zong, Yang-Wai Chow, Willy Susilo, Joonsang Baek 等USENIX Security 2025
- De-AntiFake: Rethinking the Protective Perturbations Against Voice Cloning AttacksWei Fan, Kejiang Chen, Chang Liu, Weiming Zhang 等ICML 2025
- DeAR: A Deep-Learning-Based Audio Re-recording Resilient WatermarkingChang Liu, Jie Zhang, Han Fang, Zehua Ma 等AAAI 2023 · 被引用 67 次
- XAttnMark: Learning Robust Audio Watermarking with Cross-AttentionYixin Liu, Lie Lu, Jihui Jin, Lichao Sun 等ICML 2025
