Proactive Detection of Voice Cloning with Localized Watermarking
Robin San Roman, Pierre Fernandez, Hady Elsahar, Alexandre Défossez, Teddy Furon, Tuan Tran
摘要
In the rapidly evolving field of speech generative models, there is a pressing need to ensure audio authenticity against the risks of voice cloning. We present AudioSeal, the first audio watermarking technique designed specifically for localized detection of AI-generated speech. Au-dioSeal employs a generator / detector architecture trained jointly with a localization loss to enable localized watermark detection up to the sample level, and a novel perceptual loss inspired by auditory masking, that enables Au-dioSeal to achieve better imperceptibility. Au-dioSeal achieves state-of-the-art performance in terms of robustness to real life audio manipulations and imperceptibility based on automatic and human evaluation metrics. Additionally, Au-dioSeal is designed with a fast, single-pass detector, that significantly surpasses existing models in speed, achieving detection up to two orders of magnitude faster, making it ideal for large-scale and real-time applications. Code is available at github.com/facebookresearch/audioseal.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- Can Simple Averaging Defeat Modern Watermarks?Pei Yang, Hai Ci, Yiren Song, Mike Zheng ShouNeurIPS 2024 · 被引用 48 次
- GS-Hider: Hiding Messages into 3D Gaussian SplattingXuanyu Zhang, Jiarui Meng, Runyi Li, Zhipei Xu 等NeurIPS 2024 · 被引用 43 次
- Images that Sound: Composing Images and Sounds on a Single CanvasZiyang Chen, Daniel Geng, Andrew OwensNeurIPS 2024 · 被引用 22 次
- Watermarking Autoregressive Image GenerationNikola Jovanovic, Ismail Labiad, Tomás Soucek, Martin T. Vechev 等NeurIPS 2025 · 被引用 21 次
- GROOT: Generating Robust Watermark for Diffusion-Model-Based Audio SynthesisWeizhi Liu, Yue Li, Dongdong Lin, Hui Tian 等ACM MM 2024 · 被引用 12 次
它引用的顶会 Paper16
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 被引用 2,890 次
- Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-SpeechJaehyeon Kim, Jungil Kong, Juhee SonICML 2021 · 被引用 1,267 次
- High-Fidelity Audio Compression with Improved RVQGANRithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar 等NeurIPS 2023 · 被引用 910 次
- A Watermark for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz 等ICML 2023 · 被引用 854 次
- Simple and Controllable Music GenerationJade Copet, Felix Kreuk, Itai Gat, Tal Remez 等NeurIPS 2023 · 被引用 843 次
相关 Paper
- XAttnMark: Learning Robust Audio Watermarking with Cross-AttentionYixin Liu, Lie Lu, Jihui Jin, Lichao Sun 等ICML 2025
- Learning to Watermark in the Latent Space of Generative ModelsSylvestre-Alvise Rebuffi, Tuan Tran, Valeriu Lacatusu, Pierre Fernandez 等ICML 2026 · 被引用 1 次
- AudioMarkNet: Audio Watermarking for Deepfake Speech DetectionWei Zong, Yang-Wai Chow, Willy Susilo, Joonsang Baek 等USENIX Security 2025
- Hidden in Plain Tokens: Simply Robust, Gradient-Free Watermark for Synthetic AudioGeorgios Milis, Yubin Qin, Yihan Wu, Heng HuangICML 2026 · 被引用 2 次
- Watermark-based Attribution of AI-Generated ContentZhengyuan Jiang, Moyang Guo, Yuepeng Hu, Yupu Wang 等ICLR 2026 · 被引用 11 次
