Proof-of-Learning: Definitions and Practice
Hengrui Jia, Mohammad Yaghini, Christopher A. Choquette-Choo, Natalie Dullerud, Anvith Thudi, Varun Chandrasekaran, Nicolas Papernot
Abstract
Training machine learning (ML) models typically involves expensive iterative optimization. Once the model’s final parameters are released, there is currently no mechanism for the entity which trained the model to prove that these parameters were indeed the result of this optimization procedure. Such a mechanism would support security of ML applications in several ways. For instance, it would simplify ownership resolution when multiple parties contest ownership of a specific model. It would also facilitate the distributed training across untrusted workers where Byzantine workers might otherwise mount a denial-ofservice by returning incorrect model updates.In this paper, we remediate this problem by introducing the concept of proof-of-learning in ML. Inspired by research on both proof-of-work and verified computations, we observe how a seminal training algorithm, stochastic gradient descent, accumulates secret information due to its stochasticity. This produces a natural construction for a proof-of-learning which demonstrates that a party has expended the compute require to obtain a set of model parameters correctly. In particular, our analyses and experiments show that an adversary seeking to illegitimately manufacture a proof-of-learning needs to perform at least as much work than is needed for gradient descent itself.We also instantiate a concrete proof-of-learning mechanism in both of the scenarios described above. In model ownership resolution, it protects the intellectual property of models released publicly. In distributed training, it preserves availability of the training procedure. Our empirical evaluation validates that our proof-of-learning mechanism is robust to variance induced by the hardware (e.g., ML accelerators) and software stacks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c008f596-3d02-417f-bc11-271f0b7785e5Cited by top-tier papers29
- The Alignment Problem from a Deep Learning PerspectiveRichard Ngo, Lawrence Chan, Sören MindermannICLR 2024 · 296 citations
- Handcrafted Backdoors in Deep Neural NetworksSanghyun Hong, Nicholas Carlini, Alexey KurakinNeurIPS 2022 · 105 citations
- Fingerprinting Deep Neural Networks Globally via Universal Adversarial PerturbationsZirui Peng, Shaofeng Li, Guoxing Chen, Cheng Zhang et al.CVPR 2022 · 66 citations
- Dataset Inference for Self-Supervised ModelsAdam Dziedzic, Haonan Duan, Muhammad Ahmad Kaleem, Nikita Dhawan et al.NeurIPS 2022 · 59 citations
- HuRef: HUman-REadable Fingerprint for Large Language ModelsBoyi Zeng, Lizheng Wang, Yuncong Hu, Yi Xu et al.NeurIPS 2024 · 48 citations
Builds on8
- SecureML: A System for Scalable Privacy-Preserving Machine LearningPayman Mohassel, Yupeng ZhangS&P 2017 · 2,107 citations
- Stealing Machine Learning Models via Prediction APIsFlorian Tramèr, Fan Zhang, Ari Juels, Michael K. Reiter et al.USENIX Security 2016 · 2,088 citations
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li et al.S&P 2019 · 1,801 citations
- Turning Your Weakness Into a Strength: Watermarking Deep Neural Networks by BackdooringYossi Adi, Carsten Baum, Moustapha Cissé, Benny Pinkas et al.USENIX Security 2018 · 832 citations
- CSI NN: Reverse Engineering of Neural Network Architectures Through Electromagnetic Side ChannelLejla Batina, Shivam Bhasin, Dirmanto Jap, Stjepan PicekUSENIX Security 2019 · 334 citations
Related papers
- Provenance of Training without Training Data: Towards Privacy-Preserving DNN Model Ownership VerificationYunpeng Liu, Kexin Li, Zhuotao Liu, Bihan Wen et al.WWW 2023 · 13 citations
- "Adversarial Examples" for Proof-of-LearningRui Zhang, Jian Liu, Yuan Ding, Zhibo Wang et al.S&P 2022 · 41 citations
- ZKROWNN: Zero Knowledge Right of Ownership for Neural NetworksNojan Sheybani, Zahra Ghodsi, Ritvik Kapila, Farinaz KoushanfarDAC 2023 · 10 citations
- Increasing the Cost of Model Extraction with Calibrated Proof of WorkAdam Dziedzic, Muhammad Ahmad Kaleem, Yu Shen Lu, Nicolas PapernotICLR 2022 · 37 citations
- Experimenting with Zero-Knowledge Proofs of TrainingSanjam Garg, Aarushi Goel, Somesh Jha, Saeed Mahloujifar et al.CCS 2023 · 31 citations
