Optimistic Verifiable Training by Controlling Hardware Nondeterminism
Megha Srivastava, Simran Arora, Dan Boneh
Abstract
The increasing compute demands of AI systems have led to the emergence of services that train models on behalf of clients lacking necessary resources. However, ensuring correctness of training and guarding against potential training-time attacks, such as data poisoning and backdoors, poses challenges. Existing works on verifiable training largely fall into two classes: proof-based systems, which are difficult to scale, and ``optimistic'' methods that consider a third-party auditor who can replicate the training process and contest the trainer. A key challenge with the latter is that nondeterminism between GPU types during training prevents exact replication of the training process, resulting in schemes that are non-robust. We propose a method that combines training in a higher precision than the target, rounding after intermediate computations, and sharing rounding decisions based on an adaptive thresholding procedure, to successfully control for nondeterminism. Across three different NVIDIA GPUs (A40, Titan XP, RTX 2080 Ti), we achieve exact training replication at FP32 precision for both full-training and fine-tuning of ResNet-50 (23M) and GPT-2 (117M) models. Our verifiable training scheme significantly decreases the storage and time costs compared to proof-based systems, and is publicly released at https://github.com/meghabyte/verifiable-training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a7be764f-6281-4c8d-b6c0-0594be6f1dfaCited by top-tier papers5
- User-side Model Consistency Monitoring for Open Source Large Language Models Inference ServicesQijun Miao, Zhixuan FangACL 2025 · 1 citation
- Founding Zero-Knowledge Proof of Training on Optimum VicinityGefei Tan, Adrià Gascón, Sarah Meiklejohn, Mariana Raykova et al.CCS 2025 · 1 citation
- TAO: Tolerance-Aware Optimistic Verification for Floating-Point Neural NetworksJianzhu Yao, Hongxu Su, Taobo Liao, Zerui Cheng et al.EuroSys 2026
- CloserToMe: A Unified Framework for Accurate and Transferable Latency Prediction Across Heterogeneous DevicesCheng Tang, Guochong Sui, Wenqi Lou, Zihan Wang et al.AAAI 2026
- Zero-Knowledge Location Privacy via Accurate Floating-Point SNARKsJens Ernstberger, Chengru Zhang, Luca Ciprian, Philipp Jovanovic et al.S&P 2025
Builds on7
- Poisoning Language Models During Instruction TuningAlexander Wan, Eric Wallace, Sheng Shen, Dan KleinICML 2023 · 319 citations
- Proof-of-Learning: Definitions and PracticeHengrui Jia, Mohammad Yaghini, Christopher A. Choquette-Choo, Natalie Dullerud et al.S&P 2021 · 132 citations
- Tools for Verifying Neural Models' Training DataDami Choi, Yonadav Shavit, David Kristjanson DuvenaudNeurIPS 2023 · 36 citations
- Experimenting with Zero-Knowledge Proofs of TrainingSanjam Garg, Aarushi Goel, Somesh Jha, Saeed Mahloujifar et al.CCS 2023 · 31 citations
- Zero-Knowledge Proofs of Training for Deep Neural NetworksKasra Abbaszadeh, Christodoulos Pappas, Jonathan Katz, Dimitrios PapadopoulosCCS 2024 · 24 citations
Related papers
- TrainVerify: Equivalence-Based Verification for Distributed LLM TrainingYunchi Lu, Youshan Miao, Cheng Tan, Peng Huang et al.SOSP 2025 · 1 citation
- Provable Defense Against Geometric TransformationsRem Yang, Jacob Laurel, Sasa Misailovic, Gagandeep SinghICLR 2023 · 1 citation
- Identifying and Mitigating Errors in Gradient Aggregation of Distributed Data Parallel TrainingZhenheng Tang, Junlin Huang, Zichen TANG, Xueze Kang et al.ICML 2026
- SpareTrain: Fault-Tolerant LLM Training via Low-Cost Dual Modular RedundancyRihae Park, Yeonjae Kim, Seung Yul Lee, Yeonhong Park et al.ICLR 2026
- Fast Adversarial Training with Dynamic Batch-level Attack ControlJaewon Jung, Jaeyong Song, Hongsun Jang, Hyeyoon Lee et al.DAC 2023 · 2 citations
