Tools for Verifying Neural Models' Training Data
Dami Choi, Yonadav Shavit, David Kristjanson Duvenaud
摘要
It is important that consumers and regulators can verify the provenance of large neural models to evaluate their capabilities and risks. We introduce the concept of a "Proof-of-Training-Data": any protocol that allows a model trainer to convince a Verifier of the training data that produced a set of model weights. Such protocols could verify the amount and kind of data and compute used to train the model, including whether it was trained on specific harmful or beneficial data sources. We explore efficient verification strategies for Proof-of-Training-Data that are compatible with most current large-model training procedures. These include a method for the model-trainer to verifiably pre-commit to a random seed used in training, and a method that exploits models' tendency to temporarily overfit to training data in order to detect whether a given data-point was included in training. We show experimentally that our verification procedures can catch a wide variety of attacks, including all known attacks from the Proof-of-Learning literature. Preprint. Under review. *First two authors contributed equally.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Rethinking LLM Memorization through the Lens of Adversarial CompressionAvi Schwarzschild, Zhili Feng, Pratyush Maini, Zachary C. Lipton 等NeurIPS 2024 · 被引用 120 次
- Optimistic Verifiable Training by Controlling Hardware NondeterminismMegha Srivastava, Simran Arora, Dan BonehNeurIPS 2024 · 被引用 14 次
- Holding Secrets Accountable: Auditing Privacy-Preserving Machine LearningHidde Lycklama, Alexander Viand, Nicolas Küchler, Christian Knabenhans 等USENIX Security 2024 · 被引用 11 次
- Founding Zero-Knowledge Proof of Training on Optimum VicinityGefei Tan, Adrià Gascón, Sarah Meiklejohn, Mariana Raykova 等CCS 2025 · 被引用 1 次
- TAO: Tolerance-Aware Optimistic Verification for Floating-Point Neural NetworksJianzhu Yao, Hongxu Su, Taobo Liao, Zerui Cheng 等EuroSys 2026
它引用的顶会 Paper9
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Linear Mode Connectivity and the Lottery Ticket HypothesisJonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICML 2020 · 被引用 750 次
- What Neural Networks Memorize and Why: Discovering the Long Tail via Influence EstimationVitaly Feldman, Chiyuan ZhangNeurIPS 2020 · 被引用 674 次
- Detecting AI Trojans Using Meta Neural AnalysisXiaojun Xu, Qi Wang, Huichen Li, Nikita Borisov 等S&P 2021 · 被引用 381 次
- Manipulating SGD with Data Ordering AttacksIlia Shumailov, Zakhar Shumaylov, Dmitry Kazhdan, Yiren Zhao 等NeurIPS 2021 · 被引用 125 次
相关 Paper
- Provenance of Training without Training Data: Towards Privacy-Preserving DNN Model Ownership VerificationYunpeng Liu, Kexin Li, Zhuotao Liu, Bihan Wen 等WWW 2023 · 被引用 13 次
- Proof-of-Learning: Definitions and PracticeHengrui Jia, Mohammad Yaghini, Christopher A. Choquette-Choo, Natalie Dullerud 等S&P 2021 · 被引用 132 次
- Training Data Provenance Verification: Did Your Model Use Synthetic Data from My Generative Model for Training?Yuechen Xie, Jie Song, Huiqiong Wang, Mingli SongCVPR 2025
- "Adversarial Examples" for Proof-of-LearningRui Zhang, Jian Liu, Yuan Ding, Zhibo Wang 等S&P 2022 · 被引用 41 次
- Verification of Machine Unlearning is FragileBinchi Zhang, Zihan Chen, Cong Shen, Jundong LiICML 2024 · 被引用 21 次
