LLM-Driven Treatment Effect Estimation Under Inference Time Text Confounding
Yuchen Ma, Dennis Frauen, Jonas Schweisthal, Stefan Feuerriegel
Abstract
Estimating treatment effects is crucial for personalized decision-making in medicine, but this task faces unique challenges in clinical practice. At training time, models for estimating treatment effects are typically trained on well-structured medical datasets that contain detailed patient information. However, at inference time, predictions are often made using textual descriptions (e.g., descriptions with self-reported symptoms), which are incomplete representations of the original patient information. In this work, we make three contributions. (1) We show that the discrepancy between the data available during training time and inference time can lead to biased estimates of treatment effects. We formalize this issue as an inference time text confounding problem, where confounders are fully observed during training time but only partially available through text at inference time. (2) To address this problem, we propose a novel framework for estimating treatment effects that explicitly accounts for inference time text confounding. Our framework leverages large language models (LLMs) together with a custom doubly robust learner to mitigate biases caused by the inference time text confounding. (3) Through a series of experiments, we demonstrate the effectiveness of our framework in real-world applications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d7cf0f03-3259-4845-bac4-cdcc999b9f7dCited by top-tier papers1
Ask how each one uses itBuilds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- MDAgents: An Adaptive Collaboration of LLMs for Medical Decision-MakingYubin Kim, Chanwoo Park, Hyewon Jeong, Yik Siu Chan et al.NeurIPS 2024 · 291 citations
- Can Large Language Models Infer Causation from Correlation?Zhijing Jin, Jiarui Liu, Zhiheng Lyu, Spencer Poff et al.ICLR 2024 · 186 citations
- On Inductive Biases for Heterogeneous Treatment Effect EstimationAlicia Curth, Mihaela van der SchaarNeurIPS 2021 · 114 citations
- Foundation Models for Causal Inference via Prior-Data Fitted NetworksYuchen Ma, Dennis Frauen, Emil Javurek, Stefan FeuerriegelICLR 2026 · 37 citations
Related papers
- End-To-End Causal Effect Estimation from Unstructured Natural Language DataNikita Dhawan, Leonardo Cotta, Karen Ullrich, Rahul G. Krishnan et al.NeurIPS 2024 · 24 citations
- Causal Transformer for Estimating Counterfactual OutcomesValentyn Melnychuk, Dennis Frauen, Stefan FeuerriegelICML 2022 · 146 citations
- Proximal Causal Inference With Text DataJacob M. Chen, Rohit Bhattacharya, Katherine A. KeithNeurIPS 2024 · 10 citations
- Time Series Deconfounder: Estimating Treatment Effects over Time in the Presence of Hidden ConfoundersIoana Bica, Ahmed M. Alaa, Mihaela van der SchaarICML 2020 · 133 citations
- Traceable Latent Variable Discovery Based on Multi-Agent CollaborationHuaming Du, Tao Hu, Yijie Huang, Yu Zhao et al.WWW 2026
