TAO: Tolerance-Aware Optimistic Verification for Floating-Point Neural Networks
Jianzhu Yao, Hongxu Su, Taobo Liao, Zerui Cheng, Huan Zhang, Xuechao Wang, Pramod Viswanath
摘要
Neural networks increasingly run on hardware outside the user's control (cloud GPUs, inference marketplaces, edge specialized accelerators) for both training and inference. Yet ML-as-a-Service reveals little about what actually ran or whether returned outputs faithfully reflect the intended inputs and models. Users lack recourse against service downgrades such as model swaps, quantization, graph rewrites, or discrepancies like altered advertisement embeddings. Verifying outputs is especially difficult because floating-point execution on heterogeneous accelerators is inherently non-deterministic. Existing approaches like zkML, deterministic replay, TEEs, and replication are either impractical for real floating-point neural networks or reintroduce the need to trust the vendor. We present TAO: a Tolerance-Aware Optimistic verification protocol for floating-point neural networks that accepts outputs within principled operator-level acceptance regions rather than requiring bitwise equality. TAO combines two complementary error models: (i) sound per-operator IEEE-754 worst-case bounds and (ii) tight empirical percentile profiles calibrated across hardware types. Discrepancies are resolved via a Merkle-anchored, threshold-guided interactive dispute game that recursively partitions the traced computation graph until one operator remains; at the leaf, adjudication reduces to either a lightweight theoretical-bound check or a small honest-majority vote against empirical thresholds. Unchallenged results finalize after a challenge window, without requiring trusted hardware or deterministic kernels.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph CompilationJason Ansel, Edward Z. Yang, Horace He, Natalia Gimelshein 等ASPLOS 2024 · 被引用 693 次
- Arbitrum: Scalable, private smart contractsHarry A. Kalodner, Steven Goldfeder, Xiaoqi Chen, S. Matthew Weinberg 等USENIX Security 2018 · 被引用 353 次
- Proof-of-Learning: Definitions and PracticeHengrui Jia, Mohammad Yaghini, Christopher A. Choquette-Choo, Natalie Dullerud 等S&P 2021 · 被引用 132 次
- ZKML: An Optimizing System for ML Inference in Zero-Knowledge ProofsBing-Jyue Chen, Suppakit Waiwitlikhit, Ion Stoica, Daniel KangEuroSys 2024 · 被引用 65 次
相关 Paper
- No Soundness in the Real World: On the Challenges of the Verification of Deployed Neural NetworksAttila Szász, Balázs Bánhelyi, Márk JelasityICML 2025
- Optimistic Verifiable Training by Controlling Hardware NondeterminismMegha Srivastava, Simran Arora, Dan BonehNeurIPS 2024 · 被引用 14 次
- DeepProve: Verifiable End-to-End Large Language Model InferenceNicolas Gailly, Ismael Hishon-Rezaizadeh, Tianyi Liu, Nicholas Mainardi 等CCS 2026
- Causes and Effects of Unanticipated Numerical Deviations in Neural Network Inference FrameworksAlexander Schlögl, Nora Hofer, Rainer BöhmeNeurIPS 2023 · 被引用 31 次
- CRONUS: Fault-isolated, Secure and High-performance Heterogeneous Computing for Trusted Execution EnvironmentJianyu Jiang, Ji Qi, Tianxiang Shen, Xusheng Chen 等MICRO 2022 · 被引用 27 次
