Tools for Verifying Neural Models' Training Data
Dami Choi, Yonadav Shavit, David Kristjanson Duvenaud
Abstract
It is important that consumers and regulators can verify the provenance of large neural models to evaluate their capabilities and risks. We introduce the concept of a "Proof-of-Training-Data": any protocol that allows a model trainer to convince a Verifier of the training data that produced a set of model weights. Such protocols could verify the amount and kind of data and compute used to train the model, including whether it was trained on specific harmful or beneficial data sources. We explore efficient verification strategies for Proof-of-Training-Data that are compatible with most current large-model training procedures. These include a method for the model-trainer to verifiably pre-commit to a random seed used in training, and a method that exploits models' tendency to temporarily overfit to training data in order to detect whether a given data-point was included in training. We show experimentally that our verification procedures can catch a wide variety of attacks, including all known attacks from the Proof-of-Learning literature. Preprint. Under review. *First two authors contributed equally.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 541d7d2b-9be5-4ee2-9b27-21fc879924f2Cited by top-tier papers7
- Rethinking LLM Memorization through the Lens of Adversarial CompressionAvi Schwarzschild, Zhili Feng, Pratyush Maini, Zachary C. Lipton et al.NeurIPS 2024 · 120 citations
- Optimistic Verifiable Training by Controlling Hardware NondeterminismMegha Srivastava, Simran Arora, Dan BonehNeurIPS 2024 · 14 citations
- Holding Secrets Accountable: Auditing Privacy-Preserving Machine LearningHidde Lycklama, Alexander Viand, Nicolas Küchler, Christian Knabenhans et al.USENIX Security 2024 · 11 citations
- Founding Zero-Knowledge Proof of Training on Optimum VicinityGefei Tan, Adrià Gascón, Sarah Meiklejohn, Mariana Raykova et al.CCS 2025 · 1 citation
- TAO: Tolerance-Aware Optimistic Verification for Floating-Point Neural NetworksJianzhu Yao, Hongxu Su, Taobo Liao, Zerui Cheng et al.EuroSys 2026
Builds on9
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Linear Mode Connectivity and the Lottery Ticket HypothesisJonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICML 2020 · 750 citations
- What Neural Networks Memorize and Why: Discovering the Long Tail via Influence EstimationVitaly Feldman, Chiyuan ZhangNeurIPS 2020 · 674 citations
- Detecting AI Trojans Using Meta Neural AnalysisXiaojun Xu, Qi Wang, Huichen Li, Nikita Borisov et al.S&P 2021 · 381 citations
- Manipulating SGD with Data Ordering AttacksIlia Shumailov, Zakhar Shumaylov, Dmitry Kazhdan, Yiren Zhao et al.NeurIPS 2021 · 125 citations
Related papers
- Provenance of Training without Training Data: Towards Privacy-Preserving DNN Model Ownership VerificationYunpeng Liu, Kexin Li, Zhuotao Liu, Bihan Wen et al.WWW 2023 · 13 citations
- Proof-of-Learning: Definitions and PracticeHengrui Jia, Mohammad Yaghini, Christopher A. Choquette-Choo, Natalie Dullerud et al.S&P 2021 · 132 citations
- Training Data Provenance Verification: Did Your Model Use Synthetic Data from My Generative Model for Training?Yuechen Xie, Jie Song, Huiqiong Wang, Mingli SongCVPR 2025
- "Adversarial Examples" for Proof-of-LearningRui Zhang, Jian Liu, Yuan Ding, Zhibo Wang et al.S&P 2022 · 41 citations
- Verification of Machine Unlearning is FragileBinchi Zhang, Zihan Chen, Cong Shen, Jundong LiICML 2024 · 21 citations
