TabWak: A Watermark for Tabular Diffusion Models
Chaoyi Zhu, Jiayi Tang, Jeroen M. Galjaard, Pin-Yu Chen, Robert Birke, Cornelis Bos, Lydia Y. Chen
摘要
Synthetic data offers alternatives for data augmentation and sharing. Till date, it remains unknown how to use watermarking techniques to trace and audit synthetic tables generated by tabular diffusion models to mitigate potential misuses. In this paper, we design TabWak, the first watermarking method to embed invisible signatures that control the sampling of Gaussian latent codes used to synthesize table rows via the diffusion backbone. TabWak has two key features. Different from existing image watermarking techniques, TabWak uses self-cloning and shuffling to embed the secret key in positional information of random seeds that control the Gaussian latents, allowing to use different seeds at each row for high inter-row diversity and enabling row-wise detectability. To further boost the robustness of watermark detection against post-editing attacks, TabWak uses a valid-bit mechanism that focuses on the tail of the latent code distribution for superior noise resilience. We provide theoretical guarantees on the row diversity and effectiveness of detectability. We evaluate TabWak on five datasets against baselines to show that the quality of watermarked tables remains nearly indistinguishable from nonwatermarked tables while achieving high detectability in the presence of strong post-editing attacks, with a 100% true positive rate at a 0.1% false positive rate on synthetic tables with fewer than 300 rows. Our code is available at the following repository https://github.com/chaoyitud/TabWak.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- MUSE: Model-Agnostic Tabular Watermarking via Multi-Sample SelectionLiancheng Fang, Aiwei Liu, Henry Peng Zou, Yankai Chen 等ICLR 2026 · 被引用 5 次
- TimeWak: Temporal Chained-Hashing Watermark for Time Series DataZhi Wen Soi, Chaoyi Zhu, Fouad Abiad, Aditya Shankar 等NeurIPS 2025 · 被引用 5 次
- CheckMate! Watermarking Graph Diffusion Models in Polynomial TimeRoberto Gheda, Abele Mălan, Robert Birke, Maksim Kitsak 等ICLR 2026
它引用的顶会 Paper18
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- A Watermark for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz 等ICML 2023 · 被引用 854 次
- Diffusion Models for Adversarial PurificationWeili Nie, Brandon Guo, Yujia Huang, Chaowei Xiao 等ICML 2022 · 被引用 663 次
- TabDDPM: Modelling Tabular Data with Diffusion ModelsAkim Kotelnikov, Dmitry Baranchuk, Ivan Rubachev, Artem BabenkoICML 2023 · 被引用 518 次
相关 Paper
- Guidance Watermarking for Diffusion ModelsEnoal Gesny, Eva Giboulot, Teddy Furon, Vivien ChappelierICLR 2026 · 被引用 5 次
- Attack-Resilient Image Watermarking Using Stable DiffusionLijun Zhang, Xiao Liu, Antoni Viros Martin, Cindy Xiong Bearfield 等NeurIPS 2024 · 被引用 62 次
- SEAL: Semantic Aware Image WatermarkingKasra Arabi, R. Teal Witter, Chinmay Hegde, Niv CohenICCV 2025 · 被引用 22 次
- The Stable Signature: Rooting Watermarks in Latent Diffusion ModelsPierre Fernandez, Guillaume Couairon, Hervé Jégou, Matthijs Douze 等ICCV 2023 · 被引用 370 次
- Learning to Watermark in the Latent Space of Generative ModelsSylvestre-Alvise Rebuffi, Tuan Tran, Valeriu Lacatusu, Pierre Fernandez 等ICML 2026 · 被引用 1 次
