Lune

NeurIPS2025Top-tier venue

Sample-Conditional Coverage in Split-Conformal Prediction

John C. Duchi

2025Year
2Citations
2Top-tier citations

Abstract

We revisit the problem of constructing predictive confidence sets for which we wish to obtain some type of conditional validity. We provide new arguments showing how "split conformal" methods achieve near desired coverage levels with high probability, a guarantee conditional on the validation data rather than marginal over it. In addition, we directly consider (approximate) conditional coverage, where, e.g., conditional on a covariate X belonging to some group of interest, we seek a guarantee that a predictive set covers the true outcome Y . We show that the natural method of performing quantile regression on a held-out (validation) dataset yields minimax optimal guarantees of coverage in these cases. Complementing these positive results, we also provide experimental evidence highlighting work that remains to develop computationally efficient valid predictive inference methods. provides the guarantee P (S n+1 > τ ) ≤ α. Written differently, the confidence set

39th Conference on Neural Information Processing Systems (NeurIPS 2025).

Instead of the marginal guarantee (1), we could target conditional coverage, where we say a set valued mapping C n : X ⇒ Y achieves distribution-free conditional (1 -α) coverage if for any P , when (X i , Y i ) iid ∼ P and C n is a function of (X i , Y i ) n i=1 , then for P -almost-all x,

Vovk [30] shows this is impossible. For example, when Y = R, the Lebesgue measure Leb( C(x)) is almost always infinite [30, Proposition 4] (see also extensions in [4] and [10, Corollary 7.1]): Corollary 1.1 ([30, 4, 10]). Let X be a metric space and assume X ∈ X has continuous distribution.

If C provides distribution free (1 -α) conditional coverage, then for P -almost all x ∈ X ,

These failures motivate relaxing the conditional coverage condition (2). The simplest approach considers group-conditional coverage, where for groups G ⊂ X , one targets the guarantee

Barber et al. [4, Sec. 4] achieve the coverage (3) by considering worst-case coverage over groups G; Jung et al. [17] provide variations. Gibbs, Cherian, and Candès [12] extend this idea, beginning by observing that conditional coverage P(Y ∈ C(x) | X = x) = 1 -α holds if and only if

for all bounded w. Similarly, the one-sided inequality (2) holds if and only if

for all nonnegative bounded w. Taking w(x) = 1x ∈ G for groups G ⊂ X implies the groupconditional coverage (3); relaxing the condition (4) by considering subclasses of weighting functions W ⊂ X → R leads to the following definition [12]:

Gibbs et al.'s main two examples take W of the form W = w | w(x) = ⟨v, ϕ(x)⟩ for some feature mapping ϕ : X → R d or to correspond to a reproducing kernel Hilbert space. On a new example X n+1 they perform full conformal inference [31], where implicitly for each t ∈ R, they solve

for the quantile loss ℓ α (t) = α [t] + + (1 -α) [-t] + , then define the implicit confidence set

A careful duality calculation [12, Sec. 4] shows how to compute C n by solving a linear program over O(n + d) variables using (X i ) n+1 i=1 and S n 1 = (S 1 , . . . , S n ), and Gibbs et al. show the set (5) satisfies

for w ∈ W, where ϵ int (w) is a small interpolation error term. Defining P w (A) = E P [w(X)1A]/E P [w(X)] to be the w-weighted probability of an event A for w ≥ 0, this inequality strengthens inequality (1) to imply that for all w ≥ 0, w ∈ W,

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext bd3de735-ea24-4d9f-8e1b-85f5e285a6c6

Cited by top-tier papers2

Ask how each one uses it

Builds on2

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines