Anchor Watermark: Robust Attribution for Diffusion-based Text-to-Audio Model
Xianjin Rong, Donghui Hu
Abstract
With the increasing commercialization of the latent diffusion-based text-to-audio generation, model attribution has become a critical challenge. Embedding watermarks in generated audio is an effective way to distinguish synthetic from natural audio. However, existing watermarking methods often suffer from limited robustness or require additional training, limiting their scalability in practical applications. In this paper, we propose an anchor-based inversion optimization framework. The method embeds a watermark into the model's initial latent vector, designated as a pivotal anchor, and extracts the watermark through inversion. To mitigate error accumulation and enhance robustness during inversion, we leverage the temporal consistency and distributional similarity of diffusion models, formulating watermark extraction as a time-series optimization problem. Specifically, given a suspicious audio sample and a candidate model with a predefined anchor, we first perform unguided denoising diffusion on the anchor to generate an intermediate latent trajectory as the anchor sequence. Then, we optimize the inversion process to align the inverted trajectory with the anchor sequence, thereby reducing accumulated errors. During optimization, we adopt Soft Dynamic Time Warping as the loss function. Its flexible temporal alignment capability ensures that correct attribution is achieved only when the anchor matches the target audio. Experimental results show that our method enables training-free attribution while preserving audio quality and achieving strong robustness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dbc86f1d-c851-456c-9bd4-60f288b9b3d7Builds on14
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 2,890 citations
- DPM-Solver: A Fast ODE Solver for Diffusion Probabilistic Model Sampling in Around 10 StepsCheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen et al.NeurIPS 2022 · 2,653 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
- Pseudo Numerical Methods for Diffusion Models on ManifoldsLuping Liu, Yi Ren, Zhijie Lin, Zhou ZhaoICLR 2022 · 861 citations
Related papers
- GROOT: Generating Robust Watermark for Diffusion-Model-Based Audio SynthesisWeizhi Liu, Yue Li, Dongdong Lin, Hui Tian et al.ACM MM 2024 · 12 citations
- XAttnMark: Learning Robust Audio Watermarking with Cross-AttentionYixin Liu, Lie Lu, Jihui Jin, Lichao Sun et al.ICML 2025
- RECOVER: Reliable Detection of Unauthorized Data Usage in Text-to-Image Diffusion Models via Inversion RobustnessYanhao Wei, Xiaokang Zhao, Boheng Li, Yang Zhang et al.ICML 2026
- FARI: Robust One-Step Inversion for Watermarking in Diffusion ModelsJindong Yang, Han Fang, Weiming Zhang, Nenghai Yu et al.ICLR 2026 · 1 citation
- Attention to Neural Plagiarism: Diffusion Models Can Plagiarize Your Copyrighted Images!Zihang Zou, Boqing Gong, Liqiang WangICCV 2025 · 2 citations
