Causal Estimation for Text Data with (Apparent) Overlap Violations
Lin Gui, Victor Veitch
摘要
Consider the problem of estimating the causal effect of some attribute of a text document; for example: what effect does writing a polite vs. rude email have on response time? To estimate a causal effect from observational data, we need to adjust for confounding aspects of the text that affect both the treatment and outcome-e.g., the topic or writing level of the text. These confounding aspects are unknown a priori, so it seems natural to adjust for the entirety of the text (e.g., using a transformer). However, causal identification and estimation procedures rely on the assumption of overlap: for all levels of the adjustment variables, there is randomness leftover so that every unit could have (not) received treatment. Since the treatment here is itself an attribute of the text, it is perfectly determined, and overlap is apparently violated. The purpose of this paper is to show how to handle causal identification and obtain robust causal estimation in the presence of apparent overlap violations. In brief, the idea is to use supervised representation learning to produce a data representation that preserves confounding information while eliminating information that is only predictive of the treatment. This representation then suffices for adjustment and satisfies overlap. Adapting results on non-parametric estimation, we find that this procedure is robust to conditional outcome misestimation, yielding a low-absolute-bias estimator with valid uncertainty quantification under weak conditions. Empirical results show strong improvements in bias and uncertainty quantification relative to the natural baseline. Code, demo data and a tutorial are available at https://github.com/gl-ybnbxb/TI-estimator .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Controlling Learned Effects to Reduce Spurious Correlations in Text ClassifiersParikshit Bansal, Amit SharmaACL 2023 · 被引用 1 次
- RATE: Causal Explainability of Reward Models with Imperfect CounterfactualsDavid Reber, Sean M. Richardson, Todd Nief, Cristina Garbacea 等ICML 2025
- Preference Learning for AI Alignment: a Causal PerspectiveKasia Kobalczyk, Mihaela van der SchaarICML 2025
- A Reality Check on Context Utilisation for Retrieval-Augmented GenerationLovisa Hagström, Sara Vera Marjanovic, Haeun Yu, Arnav Arora 等ACL 2025
它引用的顶会 Paper2
- Quantifying the Causal Effects of Conversational TendenciesJustine Zhang, Sendhil Mullainathan, Cristian Danescu-Niculescu-MizilCSCW 2020 · 被引用 24 次
- Text and Causal Inference: A Review of Using Text to Remove Confounding from Causal EstimatesKatherine A. Keith, David D. Jensen, Brendan O'ConnorACL 2020 · 被引用 16 次
相关 Paper
- Bounds on Representation-Induced Confounding Bias for Treatment Effect EstimationValentyn Melnychuk, Dennis Frauen, Stefan FeuerriegelICLR 2024 · 被引用 23 次
- Bayesian Topic Regression for Causal InferenceMaximilian Ahrens, Julian Ashwin, Jan-Peter Calliess, Vu NguyenEMNLP 2021
- Conditional Instrumental Variable Regression with Representation Learning for Causal InferenceDebo Cheng, Ziqi Xu, Jiuyong Li, Lin Liu 等ICLR 2024 · 被引用 14 次
- Quantifying Ignorance in Individual-Level Causal-Effect Estimates under Hidden ConfoundingAndrew Jesson, Sören Mindermann, Yarin Gal, Uri ShalitICML 2021 · 被引用 66 次
- Learning Linear Causal Representations from General Environments: Identifiability and Intrinsic AmbiguityJikai Jin, Vasilis SyrgkanisNeurIPS 2024 · 被引用 10 次
