Transferable Black-Box One-Shot Forging of Watermarks via Image Preference Models
Tomás Soucek, Sylvestre-Alvise Rebuffi, Pierre Fernandez, Nikola Jovanovic, Hady Elsahar, Valeriu Lacatusu, Tuan Tran, Alexandre Mourachko
Abstract
Recent years have seen a surge in interest in digital content watermarking techniques, driven by the proliferation of generative models and increased legal pressure. With an ever-growing percentage of AI-generated content available online, watermarking plays an increasingly important role in ensuring content authenticity and attribution at scale. There have been many works assessing the robustness of watermarking to removal attacks, yet, watermark forging, the scenario when a watermark is stolen from genuine content and applied to malicious content, remains underexplored. In this work, we investigate watermark forging in the context of widely used post-hoc image watermarking. Our contributions are as follows. First, we introduce a preference model to assess whether an image is watermarked. The model is trained using a ranking loss on purely procedurally generated images without any need for real watermarks. Second, we demonstrate the model's capability to remove and forge watermarks by optimizing the input image through backpropagation. This technique requires only a single watermarked image and works without knowledge of the watermarking model, making our attack much simpler and more practical than attacks introduced in related work. Third, we evaluate our proposed method on a variety of post-hoc image watermarking models, demonstrating that our approach can effectively forge watermarks, questioning the security of current watermarking approaches. Our code and further resources are publicly available 3 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on21
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Diffusion Models for Adversarial PurificationWeili Nie, Brandon Guo, Yujia Huang, Chaowei Xiao et al.ICML 2022 · 663 citations
Related papers
- WMCopier: Forging Invisible Watermarks on Arbitrary ImagesZiping Dong, Chao Shuai, Zhongjie Ba, Peng Cheng et al.NeurIPS 2025 · 2 citations
- Evading Watermark based Detection of AI-Generated ContentZhengyuan Jiang, Jinghuai Zhang, Neil Zhenqiang GongCCS 2023 · 46 citations
- Black-Box Forgery Attacks on Semantic Watermarks for Diffusion ModelsAndreas Müller, Denis Lukovnikov, Jonas Thietke, Asja Fischer et al.CVPR 2025
- Hidden in the Noise: Two-Stage Robust Watermarking for ImagesKasra Arabi, Benjamin Feuer, R. Teal Witter, Chinmay Hegde et al.ICLR 2025
- Robustness of AI-Image Detectors: Fundamental Limits and Practical AttacksMehrdad Saberi, Vinu Sankar Sadasivan, Keivan Rezaei, Aounon Kumar et al.ICLR 2024 · 92 citations
