Lune

NeurIPS2025Top-tier venue

Secure and Confidential Certificates of Online Fairness

Olive Franzese, Ali Shahin Shamsabadi, Carter Luck, Hamed Haddadi

2025Year
10Citations
3Top-tier citations

Abstract

The "black-box service model" enables ML service providers to serve clients while keeping their intellectual property and client data confidential. Confidentiality is critical for delivering ML services legally and responsibly, but makes it difficult for outside parties to verify important model properties such as fairness. Existing methods that assess model fairness confidentially lack either (i) reliability because they certify fairness with respect to a static set of data, and therefore fail to guarantee fairness in the presence of distribution shift or service provider malfeasance; and/or (ii) scalability due to the computational overhead of confidentiality-preserving cryptographic primitives. We address these problems by introducing online fairness certificates, which verify that a model is fair with respect to data received by the service provider online during deployment. We then present OATH, a deployably efficient and scalable zero-knowledge proof protocol for confidential online group fairness certification. OATH exploits statistical properties of group fairness via a "cut-and-choose" style protocol, enabling scalability improvements over baselines.

• Service Phase. Clients query the service provider's model, and send the auditor cryptographic commitments to their results. They perform no expensive ZKP operations, offloading them to the service provider and auditor in the next phase.

• Audit Phase. The service provider commits to a measurement of the fairness metric across all queries in the evaluation set. Then, they verify validity on a randomly sampled subset of the queries. This "cut-and-choose"-style [53] verification provides a statistical bound on the group fairness that is robust to arbitrary malicious behavior from the service provider, while maintaining excellent scaling for large numbers of client queries.

Contributions. We summarize our contributions as follows.

• Online Group Fairness Certificate. OATH audits group fairness over large sets of data received from clients online, rather than assuming fairness will generalize from a static set of offline data.

• Scalability. We exploit statistical properties of group fairness to audit an arbitrarily large set of queries to large neural networks using a constant-sized probabilistic sample (Theorem 4.1). On neural networks with 42.5 million parameters, OATH achieves 4.4 seconds of amortized runtime per query, of which only 0.23 seconds require the client to be online.

• Confidentiality & Reliability. Our cryptographic protocols guarantee that i) the auditor learns no information about the evaluation data or model parameters; and ii) the service provider cannot tamper with the audit's measurement of fairness, except within a very small probability and effect size that do not impact practical use.

Our code is publicly available at https://github.com/cleverhans-lab/ oath-zk-online-fairness.git.

2 Background, Preliminaries & Related Work ML Preliminaries. In this work we focus on probabilistic binary classifiers. That is, we consider models that can be represented as mappings M : X × 0, 1 k → 0, 1, where X is a feature space, and 0, 1 k for some k ∈ N is the space of random seeds. We assume that one feature of each query point q ∈ X corresponds to a binary demographic attribute 0, 1 ∈ S such as sex, race, or disability.

Fairness. The ML community has proposed various fairness definitions tailored to different philosophical assumptions and contexts. We focus on group fairness [22,12,22] which ensures statistical parity across different subgroups. For simplicity, we demonstrate verification of demographic parity [7] in the main text, and generalize to other metrics in Appendix B. In this work we are interested in verifying whether demographic parity is satisfied within a public threshold, as formalized below.

Definition 2.1. [Thresholded demographic parity] A predictor Ŷ satisfies demographic parity with respect to sensitive attribute S within a public threshold θ if:

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

Cited by top-tier papers3

Ask how each one uses it

Builds on15

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines