DIPBox: A Multi-scale Testing Framework for Tracking Dataset Regeneration
Tian Dong, Yan Meng, Shaofeng Li, Guoxing Chen, Yuling Chen, Zhen Liu, Haojin Zhu, Hao Chen
摘要
Training datasets have tremendous proprietary value and are vulnerable to unauthorized copying. Existing defenses mainly focus on tracking individual data points, but pay little attention to the threat of dataset regeneration. Through a measurement study of public tumor datasets, we identify substantial real-world partial-dataset replication, raising concerns about potential license noncompliance. To counter the challenge of tracking previously unknown adversarial regeneration, our key insight is that regeneration that preserves model utility inevitably preserves measurable signals across multiple feature scales. We categorize these dataset features into sample-, set-, and distribution-level features and design four similarity metrics to accurately identify regeneration. Based on these metrics, we develop DIPBox, which to our knowledge is the first testing framework that tracks regeneration suspects via multi-scale similarity testing across a spectrum of defender access settings, from limited to full information. We further provide a learning-theoretic analysis that justifies these multi-scale metrics and formalizes an inherent utility-divergence trade-off, implying fundamental limits on evasive regeneration. Extensive experiments on 16 vision and text base datasets, 320 regenerated datasets, and 590 derived models validate that DIPBox outperforms previous solutions while characterizing its robustness and limits under three adaptive attacks. CCS Concepts • Security and privacy → Digital rights management.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper25
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan 等CCS 2016 · 被引用 7,620 次
- EfficientNetV2: Smaller Models and Faster TrainingMingxing Tan, Quoc V. LeICML 2021 · 被引用 4,239 次
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 被引用 3,959 次
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 被引用 2,453 次
相关 Paper
- Dataset Inference: Ownership Resolution in Machine LearningPratyush Maini, Mohammad Yaghini, Nicolas PapernotICLR 2021 · 被引用 155 次
- Unlocking Post-hoc Dataset Inference with Synthetic DataBihe Zhao, Pratyush Maini, Franziska Boenisch, Adam DziedzicICML 2025
- Defending Against Universal Attacks Through Selective Feature RegenerationTejas S. Borkar, Felix Heide, Lina J. KaramCVPR 2020
- ZeroMark: Towards Dataset Ownership Verification without Disclosing WatermarkJunfeng Guo, Yiming Li, Ruibo Chen, Yihan Wu 等NeurIPS 2024 · 被引用 24 次
- Unshaken by Weak Embedding: Robust Probabilistic Watermarking for Dataset Copyright ProtectionShang Wang, Tianqing Zhu, Dayong Ye, Hua Ma 等NDSS 2026
