Lune

NeurIPS2025顶会

Secure and Confidential Certificates of Online Fairness

Olive Franzese, Ali Shahin Shamsabadi, Carter Luck, Hamed Haddadi

2025年份
10被引次数
3顶会引用

摘要

The "black-box service model" enables ML service providers to serve clients while keeping their intellectual property and client data confidential. Confidentiality is critical for delivering ML services legally and responsibly, but makes it difficult for outside parties to verify important model properties such as fairness. Existing methods that assess model fairness confidentially lack either (i) reliability because they certify fairness with respect to a static set of data, and therefore fail to guarantee fairness in the presence of distribution shift or service provider malfeasance; and/or (ii) scalability due to the computational overhead of confidentiality-preserving cryptographic primitives. We address these problems by introducing online fairness certificates, which verify that a model is fair with respect to data received by the service provider online during deployment. We then present OATH, a deployably efficient and scalable zero-knowledge proof protocol for confidential online group fairness certification. OATH exploits statistical properties of group fairness via a "cut-and-choose" style protocol, enabling scalability improvements over baselines.

• Service Phase. Clients query the service provider's model, and send the auditor cryptographic commitments to their results. They perform no expensive ZKP operations, offloading them to the service provider and auditor in the next phase.

• Audit Phase. The service provider commits to a measurement of the fairness metric across all queries in the evaluation set. Then, they verify validity on a randomly sampled subset of the queries. This "cut-and-choose"-style [53] verification provides a statistical bound on the group fairness that is robust to arbitrary malicious behavior from the service provider, while maintaining excellent scaling for large numbers of client queries.

Contributions. We summarize our contributions as follows.

• Online Group Fairness Certificate. OATH audits group fairness over large sets of data received from clients online, rather than assuming fairness will generalize from a static set of offline data.

• Scalability. We exploit statistical properties of group fairness to audit an arbitrarily large set of queries to large neural networks using a constant-sized probabilistic sample (Theorem 4.1). On neural networks with 42.5 million parameters, OATH achieves 4.4 seconds of amortized runtime per query, of which only 0.23 seconds require the client to be online.

• Confidentiality & Reliability. Our cryptographic protocols guarantee that i) the auditor learns no information about the evaluation data or model parameters; and ii) the service provider cannot tamper with the audit's measurement of fairness, except within a very small probability and effect size that do not impact practical use.

Our code is publicly available at https://github.com/cleverhans-lab/ oath-zk-online-fairness.git.

2 Background, Preliminaries & Related Work ML Preliminaries. In this work we focus on probabilistic binary classifiers. That is, we consider models that can be represented as mappings M : X × 0, 1 k → 0, 1, where X is a feature space, and 0, 1 k for some k ∈ N is the space of random seeds. We assume that one feature of each query point q ∈ X corresponds to a binary demographic attribute 0, 1 ∈ S such as sex, race, or disability.

Fairness. The ML community has proposed various fairness definitions tailored to different philosophical assumptions and contexts. We focus on group fairness [22,12,22] which ensures statistical parity across different subgroups. For simplicity, we demonstrate verification of demographic parity [7] in the main text, and generalize to other metrics in Appendix B. In this work we are interested in verifying whether demographic parity is satisfied within a public threshold, as formalized below.

Definition 2.1. [Thresholded demographic parity] A predictor Ŷ satisfies demographic parity with respect to sensitive attribute S within a public threshold θ if:

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper3

问问它们各自怎么用它

它引用的顶会 Paper15

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖