Proximal Causal Inference With Text Data
Jacob M. Chen, Rohit Bhattacharya, Katherine A. Keith
摘要
Recent text-based causal methods attempt to mitigate confounding bias by estimating proxies of confounding variables that are partially or imperfectly measured from unstructured text data. These approaches, however, assume analysts have supervised labels of the confounders given text for a subset of instances, a constraint that is sometimes infeasible due to data privacy or annotation costs. In this work, we address settings in which an important confounding variable is completely unobserved. We propose a new causal inference method that uses two instances of pre-treatment text data, infers two proxies using two zero-shot models on the separate instances, and applies these proxies in the proximal g-formula. We prove, under certain assumptions about the instances of text and accuracy of the zero-shot predictions, that our method of inferring text-based proxies satisfies identification conditions of the proximal g-formula while other seemingly reasonable proposals do not. To address untestable assumptions associated with our method and the proximal g-formula, we further propose an odds ratio falsification heuristic that flags when to proceed with downstream effect estimation using the inferred proxies. We evaluate our method in synthetic and semi-synthetic settings -- the latter with real-world clinical notes from MIMIC-III and open large language models for zero-shot prediction -- and find that our method produces estimates with low bias. We believe that this text-based design of proxies allows for the use of proximal causal inference in a wider range of scenarios, particularly those for which obtaining suitable proxies from structured data is difficult.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Towards Multimodal Time Series Anomaly Detection with Semantic Alignment and Condensed InteractionShiyan Hu, Jianxin Jin, Yang Shu, Peng Chen 等ICLR 2026 · 被引用 7 次
- LLM-Driven Treatment Effect Estimation Under Inference Time Text ConfoundingYuchen Ma, Dennis Frauen, Jonas Schweisthal, Stefan FeuerriegelNeurIPS 2025 · 被引用 7 次
- Spiked-CFR: Causal Representation Learning from LLMs via Wasserstein Projection PursuitFan Wang, Hengyu Yue, Yu Bowen, Weiming Liu 等ICML 2026
它引用的顶会 Paper8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach 等ICLR 2022 · 被引用 1,976 次
- Proximal Causal Learning with Kernels: Two-Stage Estimation and Moment RestrictionAfsaneh Mastouri, Yuchen Zhu, Limor Gultchin, Anna Korba 等ICML 2021 · 被引用 78 次
- Using Imperfect Surrogates for Downstream Inference: Design-based Supervised Learning for Social Science Applications of Large Language ModelsNaoki Egami, Musashi Hinck, Brandon M. Stewart, Hanying WeiNeurIPS 2023 · 被引用 74 次
相关 Paper
- Deep Multi-Modal Structural Equations For Causal Effect Estimation With Unstructured ProxiesShachi Deshpande, Kaiwen Wang, Dhruv Sreenivas, Zheng Li 等NeurIPS 2022 · 被引用 15 次
- End-To-End Causal Effect Estimation from Unstructured Natural Language DataNikita Dhawan, Leonardo Cotta, Karen Ullrich, Rahul G. Krishnan 等NeurIPS 2024 · 被引用 24 次
- Deep Learning Methods for Proximal Inference via Maximum Moment RestrictionBenjamin Kompa, David R. Bellamy, Thomas Kolokotrones, James M. Robins 等NeurIPS 2022 · 被引用 22 次
- Traceable Latent Variable Discovery Based on Multi-Agent CollaborationHuaming Du, Tao Hu, Yijie Huang, Yu Zhao 等WWW 2026
- Optimal Treatment Regimes for Proximal Causal LearningTao Shen, Yifan CuiNeurIPS 2023 · 被引用 11 次
