Copyright-Certified Distillation Dataset: Distilling One Million Coins into One Bitcoin with Your Private Key
Tengjun Liu, Ying Chen, Wanxuan Gu
Abstract
The rapid development of neural network dataset distillation in recent years has provided new ideas in many areas such as continuous learning, neural network architecture search and privacy preservation. Dataset distillation is a very effective method to distill large training datasets into small data, thus ensuring that the test accuracy of models trained on their synthesized small datasets matches that of models trained on the full dataset. Thus, dataset distillation itself is commercially valuable, not only for reducing training costs, but also for compressing storage costs and significantly reducing the training costs of deep learning. However, copyright protection for dataset distillation has not been proposed yet, so we propose the first method to protect intellectual property by embedding watermarks in the dataset distillation process. Our approach not only popularizes the dataset distillation technique, but also authenticates the ownership of the distilled dataset by the models trained on that distilled dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ed9fc278-9f36-4297-9247-92ef0522a12eCited by top-tier papers1
Ask how each one uses itBuilds on13
- Training Generative Adversarial Networks with Limited DataTero Karras, Miika Aittala, Janne Hellsten, Samuli Laine et al.NeurIPS 2020 · 2,345 citations
- Turning Your Weakness Into a Strength: Watermarking Deep Neural Networks by BackdooringYossi Adi, Carsten Baum, Moustapha Cissé, Benny Pinkas et al.USENIX Security 2018 · 832 citations
- Differentiable Augmentation for Data-Efficient GAN TrainingShengyu Zhao, Zhijian Liu, Ji Lin, Jun-Yan Zhu et al.NeurIPS 2020 · 707 citations
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 684 citations
- Property Inference Attacks on Fully Connected Neural Networks using Permutation Invariant RepresentationsKaran Ganju, Qi Wang, Wei Yang, Carl A. Gunter et al.CCS 2018 · 574 citations
Related papers
- Watermarking Deep Neural Networks with Greedy ResidualsHanwen Liu, Zhenyu Weng, Yuesheng ZhuICML 2021 · 69 citations
- Identification for Deep Neural Network: Simply Adjusting Few Weights!Yingjie Lao, Peng Yang, Weijie Zhao, Ping LiICDE 2022 · 19 citations
- Safe Distillation BoxJingwen Ye, Yining Mao, Jie Song, Xinchao Wang et al.AAAI 2022 · 14 citations
- Free Fine-tuning: A Plug-and-Play Watermarking Scheme for Deep Neural NetworksRun Wang, Jixing Ren, Boheng Li, Tianyi She et al.ACM MM 2023 · 20 citations
- ClearStamp: A Human-Visible and Robust Model-Ownership Proof based on Transposed Model TrainingTorsten Krauß, Jasper Stang, Alexandra DmitrienkoUSENIX Security 2024 · 9 citations
