TabularMark: Watermarking Tabular Datasets for Machine Learning
Yihao Zheng, Haocheng Xia, Junyuan Pang, Jinfei Liu, Kui Ren, Lingyang Chu, Yang Cao, Li Xiong
摘要
Watermarking is broadly utilized to protect ownership of shared data while preserving data utility. However, existing watermarking methods for tabular datasets fall short on the desired properties (detectability, non-intrusiveness, and robustness) and only preserve data utility from the perspective of data statistics, ignoring the performance of downstream ML models trained on the datasets. Can we watermark tabular datasets without significantly compromising their utility for training ML models while preventing attackers from training usable ML models on attacked datasets? In this paper, we propose a hypothesis testing-based watermarking scheme, TabularMark. Data noise partitioning is utilized for data perturbation during embedding, which is adaptable for numerical and categorical attributes while preserving the data utility. For detection, a custom-threshold one proportion z-test is employed , which can reliably determine the presence of the watermark. Experiments on real-world and synthetic datasets demonstrate the superiority of TabularMark in detectability, non-intrusiveness, and robustness.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- MUSE: Model-Agnostic Tabular Watermarking via Multi-Sample SelectionLiancheng Fang, Aiwei Liu, Henry Peng Zou, Yankai Chen 等ICLR 2026 · 被引用 5 次
- ArtistAuditor: Auditing Artist Style Pirate in Text-to-Image Generation ModelsLinkang Du, Zheng Zhu, Min Chen, Zhou Su 等WWW 2025 · 被引用 5 次
- TabWak: A Watermark for Tabular Diffusion ModelsChaoyi Zhu, Jiayi Tang, Jeroen M. Galjaard, Pin-Yu Chen 等ICLR 2025
- Vector Database WatermarkingZhiwen Ren, Wei Fan, Qiyi Yao, Jing Qiu 等NeurIPS 2025
相关 Paper
- B2Mark: A Blind and Buyer-Traceable Watermarking Scheme for Tabular DatasetsYihao Zheng, Jinfei Liu, Kui Ren, Li XiongSIGMOD 2026 · 被引用 1 次
- Provable Watermarking for Data Poisoning AttacksYifan Zhu, Lijia Yu, Xiao-Shan GaoNeurIPS 2025 · 被引用 3 次
- PuzzleMark: Implicit Jigsaw Learning for Robust Code Dataset Watermarking in Neural Code Completion ModelsHaocheng Huang, Yuchen Chen, Weisong Sun, Peizhuo Lv 等FSE 2026
- ZeroMark: Towards Dataset Ownership Verification without Disclosing WatermarkJunfeng Guo, Yiming Li, Ruibo Chen, Yihan Wu 等NeurIPS 2024 · 被引用 24 次
- FreqyWM: Frequency Watermarking for the New Data EconomyDevris Isler, Elisa Cabana, Álvaro García-Recuero, Georgia Koutrika 等ICDE 2024 · 被引用 2 次
