Estimate Level Adjustment For Inference With Proxies Under Random Distribution Shifts
Steven Wilkins-Reeves, Alexandra N. M. Darmon, Deeksha Sinha
Abstract
In many scientific domains, including experimentation, researchers rely on measurements of proxy outcomes to achieve faster and more frequent reads, especially when the primary outcome of interest is challenging to measure directly. While proxies offer a more readily accessible observation for inference, the ultimate goal is to draw statistical inferences about the primary outcome parameter; and proxy data are typically imperfect in some ways. To correct for these imperfections, current statistical inference methods often depend on strict identifying assumptions (such as surrogacy, covariate/label shift, or missingness assumptions). These assumptions can be difficult to validate and may be violated by various additional sources of distribution shift, potentially leading to biased parameter estimates and miscalibrated uncertainty quantification. We introduce an estimate-level framework, inspired by domain adaptation methods, to empirically calibrate proxy-based inference. This framework models the proxy–primary metric discrepancy as a random effect at the parameter level, estimating its distribution from aggregated historical observations across past domains (e.g., experiments, time periods, or distinct segments). This method avoids the requirement for retaining individual-level response data. Additionally, this adjustment can be layered on top of existing proxy-correction methods (such as prediction-powered inference or importance weighting) to account for additional biases not addressed by those corrections. To manage uncertainty when the number of historical domains is limited, we provide both a method-of-moments estimator and a domain bootstrap procedure. Through extensive simulations, our adjusted intervals demonstrate improved calibration and more reliable coverage across various proxy procedures and deviations from standard covariate shift mechanisms. We further validate this approach using publicly available datasets and real-world experiments. Code for replicating simulations and experiments on public datasets are available at https://www.github.com/facebookresearch/estimate-level-adjustment.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aceeff69-806e-49e1-a58d-72556fbd20e3Builds on4
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- Synthetic-powered predictive inferenceMeshi Bashari, Roy Maor Lotan, Yonghoon Lee, Edgar Dobriban et al.NeurIPS 2025 · 12 citations
- Inferring the Long-Term Causal Effects of Long-Term Treatments from Short-Term ExperimentsAllen Tran, Aurélien Bibaut, Nathan KallusICML 2024 · 11 citations
- Learning the Covariance of Treatment Effects Across Many Weak ExperimentsAurélien Bibaut, Winston Chou, Simon Ejdemyr, Nathan KallusKDD 2024 · 2 citations
Related papers
- Using Surrogates in Covariate-adjusted Response-adaptive Randomization Experiments with Delayed OutcomesLei Shi, Waverly Wei, Jingshen WangNeurIPS 2024 · 4 citations
- Diffusion-Based Probabilistic Uncertainty Estimation for Active Domain AdaptationZhekai Du, Jingjing LiNeurIPS 2023 · 32 citations
- PAC Prediction Sets Under Label ShiftWenwen Si, Sangdon Park, Insup Lee, Edgar Dobriban et al.ICLR 2024 · 15 citations
- MEC: Machine-Learning-Assisted Generalized Entropy Calibration for Semi-Supervised Mean EstimationSe Yoon Lee, Jae-kwang KimICML 2026 · 1 citation
- IW-GAE: Importance weighted group accuracy estimation for improved calibration and model selection in unsupervised domain adaptationTaejong Joo, Diego KlabjanICML 2024 · 1 citation
