Speech Watermarking with Discrete Intermediate Representations
Shengpeng Ji, Ziyue Jiang, Jialong Zuo, Minghui Fang, Yifu Chen, Tao Jin, Zhou Zhao
摘要
Speech watermarking techniques can proactively mitigate the potential harmful consequences of instant voice cloning techniques. These techniques involve the insertion of signals into speech that are imperceptible to humans but can be detected by algorithms. Previous approaches typically embed watermark messages into continuous space. However, intuitively, embedding watermark information into robust discrete latent space can significantly improve the robustness of watermarking systems. In this paper, we propose DiscreteWM, a novel speech watermarking framework that injects watermarks into the discrete intermediate representations of speech. Specifically, we map speech into discrete latent space with a vector-quantized autoencoder and inject watermarks by changing the modular arithmetic relation of discrete IDs. To ensure the imperceptibility of watermarks, we also propose a manipulator model to select the candidate tokens for watermark embedding. Experimental results demonstrate that our framework achieves state-of-the-art performance in robustness and imperceptibility, simultaneously. Moreover, our flexible frame-wise approach can serve as an efficient solution for both voice cloning detection and information hiding. Additionally, DiscreteWM can encode 1 to 150 bits of watermark information within a 1-second speech clip, indicating its encoding capacity. Audio samples are available at https://DiscreteWM.github.io/discrete wm.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper14
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 被引用 2,890 次
- Voicebox: Text-Guided Multilingual Universal Speech Generation at ScaleMatthew Le, Apoorv Vyas, Bowen Shi, Brian Karrer 等NeurIPS 2023 · 被引用 613 次
- YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot Voice Conversion for EveryoneEdresson Casanova, Julian Weber, Christopher Dane Shulby, Arnaldo Cândido Júnior 等ICML 2022 · 被引用 602 次
- NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing SynthesizersKai Shen, Zeqian Ju, Xu Tan, Eric Liu 等ICLR 2024 · 被引用 362 次
- SQ-VAE: Variational Bayes on Discrete Representation with Self-annealed Stochastic QuantizationYuhta Takida, Takashi Shibuya, Wei-Hsiang Liao, Chieh-Hsin Lai 等ICML 2022 · 被引用 99 次
相关 Paper
- Detecting Voice Cloning Attacks via Timbre WatermarkingChang Liu, Jie Zhang, Tianwei Zhang, Xi Yang 等NDSS 2024
- FreqMark: Invisible Image Watermarking via Frequency Based Optimization in Latent SpaceYiyang Guo, Ruizhe Li, Mude Hui, Hanzhong Guo 等NeurIPS 2024 · 被引用 14 次
- Learning to Watermark in the Latent Space of Generative ModelsSylvestre-Alvise Rebuffi, Tuan Tran, Valeriu Lacatusu, Pierre Fernandez 等ICML 2026 · 被引用 1 次
- Robust Distortion-Free Watermark for Autoregressive Audio Generation ModelsYihan Wu, Georgios Milis, Ruibo Chen, Heng HuangNeurIPS 2025 · 被引用 5 次
- Hidden in Plain Tokens: Simply Robust, Gradient-Free Watermark for Synthetic AudioGeorgios Milis, Yubin Qin, Yihan Wu, Heng HuangICML 2026 · 被引用 2 次
