Robust Distortion-Free Watermark for Autoregressive Audio Generation Models
Yihan Wu, Georgios Milis, Ruibo Chen, Heng Huang
Abstract
The rapid advancement of next-token-prediction models has led to widespread adoption across modalities, enabling the creation of realistic synthetic media. In the audio domain, while autoregressive speech models have propelled conversational interactions forward, the potential for misuse, such as impersonation in phishing schemes or crafting misleading speech recordings, has also increased. Security measures such as watermarking have thus become essential to ensuring the authenticity of digital media. Traditional statistical watermarking methods used for autoregressive language models face challenges when applied to autoregressive audio models, due to the inevitable "retokenization mismatch" -the discrepancy between original and retokenized discrete audio token sequences. To address this, we introduce ALIGNED-IS, a novel, distortion-free watermark, specifically crafted for audio generation models. This technique utilizes a clustering approach that treats tokens within the same cluster equivalently, effectively countering the retokenization mismatch issue. Our comprehensive testing on prevalent audio generation platforms demonstrates that ALIGNED-IS not only preserves the quality of generated audio but also significantly improves the watermark detectability compared to the state-of-the-art distortion-free watermarking adaptations, establishing a new benchmark in secure audio technology applications. We release the code in https://github.com/g-milis/AlignedIS.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8a95b0a3-0f27-4e68-9661-a7fff820981cCited by top-tier papers2
- In-Context Watermarks for Large Language ModelsYepeng Liu, Xuandong Zhao, Christopher Kruegel, Dawn Song et al.ICLR 2026 · 14 citations
- Hidden in Plain Tokens: Simply Robust, Gradient-Free Watermark for Synthetic AudioGeorgios Milis, Yubin Qin, Yihan Wu, Heng HuangICML 2026 · 2 citations
Builds on16
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- A Watermark for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz et al.ICML 2023 · 854 citations
- NExT-GPT: Any-to-Any Multimodal LLMShengqiong Wu, Hao Fei, Leigang Qu, Wei Ji et al.ICML 2024 · 786 citations
- Provable Robust Watermarking for AI-Generated TextXuandong Zhao, Prabhanjan Vijendra Ananth, Lei Li, Yu-Xiang WangICLR 2024 · 312 citations
- Making LLaMA SEE and Draw with SEED TokenizerYuying Ge, Sijie Zhao, Ziyun Zeng, Yixiao Ge et al.ICLR 2024 · 202 citations
Related papers
- ClusterMark: Towards Robust Watermarking for Autoregressive Image Generators with Visual Token ClusteringDenis Lukovnikov, Andreas Müller, Erwin Quiring, Asja FischerCVPR 2026 · 3 citations
- AudioMarkNet: Audio Watermarking for Deepfake Speech DetectionWei Zong, Yang-Wai Chow, Willy Susilo, Joonsang Baek et al.USENIX Security 2025
- XAttnMark: Learning Robust Audio Watermarking with Cross-AttentionYixin Liu, Lie Lu, Jihui Jin, Lichao Sun et al.ICML 2025
- Watermarking Autoregressive Image GenerationNikola Jovanovic, Ismail Labiad, Tomás Soucek, Martin T. Vechev et al.NeurIPS 2025 · 21 citations
- Watermarking Diffusion Language ModelsThibaud Gloaguen, Robin Staab, Nikola Jovanović, Martin VechevICLR 2026 · 13 citations
