Optimistic Verifiable Training by Controlling Hardware Nondeterminism
Megha Srivastava, Simran Arora, Dan Boneh
摘要
The increasing compute demands of AI systems have led to the emergence of services that train models on behalf of clients lacking necessary resources. However, ensuring correctness of training and guarding against potential training-time attacks, such as data poisoning and backdoors, poses challenges. Existing works on verifiable training largely fall into two classes: proof-based systems, which are difficult to scale, and ``optimistic'' methods that consider a third-party auditor who can replicate the training process and contest the trainer. A key challenge with the latter is that nondeterminism between GPU types during training prevents exact replication of the training process, resulting in schemes that are non-robust. We propose a method that combines training in a higher precision than the target, rounding after intermediate computations, and sharing rounding decisions based on an adaptive thresholding procedure, to successfully control for nondeterminism. Across three different NVIDIA GPUs (A40, Titan XP, RTX 2080 Ti), we achieve exact training replication at FP32 precision for both full-training and fine-tuning of ResNet-50 (23M) and GPT-2 (117M) models. Our verifiable training scheme significantly decreases the storage and time costs compared to proof-based systems, and is publicly released at https://github.com/meghabyte/verifiable-training.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- User-side Model Consistency Monitoring for Open Source Large Language Models Inference ServicesQijun Miao, Zhixuan FangACL 2025 · 被引用 1 次
- Founding Zero-Knowledge Proof of Training on Optimum VicinityGefei Tan, Adrià Gascón, Sarah Meiklejohn, Mariana Raykova 等CCS 2025 · 被引用 1 次
- TAO: Tolerance-Aware Optimistic Verification for Floating-Point Neural NetworksJianzhu Yao, Hongxu Su, Taobo Liao, Zerui Cheng 等EuroSys 2026
- CloserToMe: A Unified Framework for Accurate and Transferable Latency Prediction Across Heterogeneous DevicesCheng Tang, Guochong Sui, Wenqi Lou, Zihan Wang 等AAAI 2026
- Zero-Knowledge Location Privacy via Accurate Floating-Point SNARKsJens Ernstberger, Chengru Zhang, Luca Ciprian, Philipp Jovanovic 等S&P 2025
它引用的顶会 Paper7
- Poisoning Language Models During Instruction TuningAlexander Wan, Eric Wallace, Sheng Shen, Dan KleinICML 2023 · 被引用 319 次
- Proof-of-Learning: Definitions and PracticeHengrui Jia, Mohammad Yaghini, Christopher A. Choquette-Choo, Natalie Dullerud 等S&P 2021 · 被引用 132 次
- Tools for Verifying Neural Models' Training DataDami Choi, Yonadav Shavit, David Kristjanson DuvenaudNeurIPS 2023 · 被引用 36 次
- Experimenting with Zero-Knowledge Proofs of TrainingSanjam Garg, Aarushi Goel, Somesh Jha, Saeed Mahloujifar 等CCS 2023 · 被引用 31 次
- Zero-Knowledge Proofs of Training for Deep Neural NetworksKasra Abbaszadeh, Christodoulos Pappas, Jonathan Katz, Dimitrios PapadopoulosCCS 2024 · 被引用 24 次
相关 Paper
- TrainVerify: Equivalence-Based Verification for Distributed LLM TrainingYunchi Lu, Youshan Miao, Cheng Tan, Peng Huang 等SOSP 2025 · 被引用 1 次
- Provable Defense Against Geometric TransformationsRem Yang, Jacob Laurel, Sasa Misailovic, Gagandeep SinghICLR 2023 · 被引用 1 次
- Identifying and Mitigating Errors in Gradient Aggregation of Distributed Data Parallel TrainingZhenheng Tang, Junlin Huang, Zichen TANG, Xueze Kang 等ICML 2026
- SpareTrain: Fault-Tolerant LLM Training via Low-Cost Dual Modular RedundancyRihae Park, Yeonjae Kim, Seung Yul Lee, Yeonhong Park 等ICLR 2026
- Fast Adversarial Training with Dynamic Batch-level Attack ControlJaewon Jung, Jaeyong Song, Hongsun Jang, Hyeyoon Lee 等DAC 2023 · 被引用 2 次
