Valid Inference with Imperfect Synthetic Data
Yewon Byun, Shantanu Gupta, Zachary C. Lipton, Rachel Leah Childers, Bryan Wilder
Abstract
Predictions and generations from large language models are increasingly being explored as an aid in limited data regimes, such as in computational social science and human subjects research. While prior technical work has mainly explored the potential to use model-predicted labels for unlabeled data in a principled manner, there is increasing interest in using large language models to generate entirely new synthetic samples (e.g., synthetic simulations), such as in responses to surveys. However, it remains unclear by what means practitioners can combine such data with real data and yet produce statistically valid conclusions upon them. In this paper, we introduce a new estimator based on generalized method of moments, providing a hyperparameter-free solution with strong theoretical guarantees to address this challenge. Intriguingly, we find that interactions between the moment residuals of synthetic data and those of real data (i.e., when they are predictive of each other) can greatly improve estimates of the target parameter. We validate the finite-sample performance of our estimator across different tasks in computational social science applications, demonstrating large empirical gains.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Doubly-Robust LLM-as-a-Judge: Externally Valid Estimation with Imperfect PersonasLuke Guerdan, Justin Whitehouse, Kimberly Truong, Ken Holstein et al.ICLR 2026 · 8 citations
- General Synthetic-Powered InferenceMeshi Bashari, Yonghoon Lee, Roy Lotan, Edgar Dobriban et al.ICML 2026 · 5 citations
- AI-Assisted Variance Reduction in Randomized ExperimentsDavid Arbour, Eli Ben-Michael, Avi Feller, Apoorva Lal et al.KDD 2026 · 4 citations
Builds on3
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris et al.UIST 2023 · 1,882 citations
- Questioning the Survey Responses of Large Language ModelsRicardo Dominguez-Olmedo, Moritz Hardt, Celestine Mendler-DünnerNeurIPS 2024 · 116 citations
- Using Imperfect Surrogates for Downstream Inference: Design-based Supervised Learning for Social Science Applications of Large Language ModelsNaoki Egami, Musashi Hinck, Brandon M. Stewart, Hanying WeiNeurIPS 2023 · 74 citations
Related papers
- Generative Augmented InferenceCheng Lu, Mengxin Wang, Dennis Zhang, Heng ZhangICML 2026 · 1 citation
- Valid Survey Simulations with Limited Human Data: The Roles of Prompting, Fine-Tuning, and RectificationStefan Krsteski, Giuseppe Russo, Serina Chang, Robert West et al.ACL 2026 · 10 citations
- What is a protest anyway? Codebook conceptualization is still a first-order concern in LLM-era classificationAndrew Halterman, Katherine A. KeithACL 2026 · 3 citations
- Uncertainty Quantification for LLM-Based Survey SimulationsChengpiao Huang, Yuhang Wu, Kaizheng WangICML 2025
- G-Sim: Generative Simulations with Large Language Models and Gradient-Free CalibrationSamuel Holt, Max Ruiz Luyten, Antonin Berthon, Mihaela van der SchaarICML 2025
