Efficient Randomized Experiments Using Foundation Models
Piersilvio De Bartolomeis, Javier Abad, Guanbo Wang, Konstantin Donhauser, Raymond M. Duch, Fanny Yang, Issa J. Dahabreh
摘要
Randomized experiments are the preferred approach for evaluating the effects of interventions, but they are costly and often yield estimates with substantial uncertainty. On the other hand, in silico experiments leveraging foundation models offer a cost-effective alternative that can potentially attain higher statistical precision. However, the benefits of in silico experiments come with a significant risk: statistical inferences are not valid if the models fail to accurately predict experimental responses to interventions. In this paper, we propose a novel approach that integrates the predictions from multiple foundation models with experimental data while preserving valid statistical inference. Our estimator is consistent and asymptotically normal, with asymptotic variance no larger than the standard estimator based on experimental data alone. Importantly, these statistical properties hold even when model predictions are arbitrarily biased. Empirical results across several randomized experiments show that our estimator offers substantial precision gains, equivalent to a reduction of up to 20% in the sample size needed to match the same precision as the standard estimator based on experimental data alone 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Prediction-Powered Causal InferencesRiccardo Cadei, Ilker Demirel, Piersilvio De Bartolomeis, Lukas Lindorfer 等NeurIPS 2025 · 被引用 9 次
- General Synthetic-Powered InferenceMeshi Bashari, Yonghoon Lee, Roy Lotan, Edgar Dobriban 等ICML 2026 · 被引用 5 次
- Revisiting Active Sequential Prediction-Powered Mean EstimationMaria-Eleni Sfyraki, Jun-Kun WangICLR 2026 · 被引用 4 次
- AI-Assisted Variance Reduction in Randomized ExperimentsDavid Arbour, Eli Ben-Michael, Avi Feller, Apoorva Lal 等KDD 2026 · 被引用 4 次
- How can we assess human-agent interactions? Case studies in software agent designValerie Chen, Rohit Malhotra, Xingyao Wang, Juan Michelini 等ICML 2026
它引用的顶会 Paper5
- Using Imperfect Surrogates for Downstream Inference: Design-based Supervised Learning for Social Science Applications of Large Language ModelsNaoki Egami, Musashi Hinck, Brandon M. Stewart, Hanying WeiNeurIPS 2023 · 被引用 74 次
- End-To-End Causal Effect Estimation from Unstructured Natural Language DataNikita Dhawan, Leonardo Cotta, Karen Ullrich, Rahul G. Krishnan 等NeurIPS 2024 · 被引用 24 次
- Prediction-powered Generalization of Causal InferencesIlker Demirel, Ahmed M. Alaa, Anthony Philippakis, David A. SontagICML 2024 · 被引用 18 次
- No Free Lunch: Non-Asymptotic Analysis of Prediction-Powered InferencePranav Mani, Peng Xu, Zachary Lipton, Michael OberstICML 2026 · 被引用 8 次
- Limits to scalable evaluation at the frontier: LLM as judge won't beat twice the dataFlorian E. Dorner, Vivian Yvonne Nastl, Moritz HardtICLR 2025
相关 Paper
- Constructing Confidence Intervals for Average Treatment Effects from Multiple DatasetsYuxin Wang, Maresa Schröder, Dennis Frauen, Jonas Schweisthal 等ICLR 2025
- Estimate Level Adjustment For Inference With Proxies Under Random Distribution ShiftsSteven Wilkins-Reeves, Alexandra N. M. Darmon, Deeksha SinhaKDD 2026 · 被引用 1 次
- Estimating Joint Treatment Effects by Combining Multiple ExperimentsYonghan Jung, Jin Tian, Elias BareinboimICML 2023 · 被引用 5 次
- Bridging Domain Expertise and Generalization for Performance EstimationShuxuan Li, Zhilin Zhao, Quyu Kong, Wei-Shi ZhengCVPR 2026 · 被引用 1 次
- Estimating Distributional Treatment Effects in Randomized Experiments: Machine Learning for Variance ReductionUndral Byambadalai, Tatsushi Oka, Shota YasuiICML 2024 · 被引用 7 次
