Bayesian Topic Regression for Causal Inference
Maximilian Ahrens, Julian Ashwin, Jan-Peter Calliess, Vu Nguyen
摘要
Causal inference using observational text data is becoming increasingly popular in many research areas. This paper presents the Bayesian Topic Regression (BTR) model that uses both text and numerical information to model an outcome variable. It allows estimation of both discrete and continuous treatment effects. Furthermore, it allows for the inclusion of additional numerical confounding factors next to text data. To this end, we combine a supervised Bayesian topic model with a Bayesian regression framework and perform supervised representation learning for the text features jointly with the regression parameter training, respecting the Frisch-Waugh-Lovell theorem. Our paper makes two main contributions. First, we provide a regression framework that allows causal inference in settings when both text and numerical confounders are of relevance. We show with synthetic and semi-synthetic datasets that our joint approach recovers ground truth with lower bias than any benchmark model, when text and numerical features are correlated. Second, experiments on two real-world datasets demonstrate that a joint and supervised learning strategy also yields superior prediction results compared to strategies that estimate regression weights for text and non-text features separately, being even competitive with more complex deep neural networks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper1
相关 Paper
- Causal Estimation for Text Data with (Apparent) Overlap ViolationsLin Gui, Victor VeitchICLR 2023 · 被引用 6 次
- LLM-Driven Treatment Effect Estimation Under Inference Time Text ConfoundingYuchen Ma, Dennis Frauen, Jonas Schweisthal, Stefan FeuerriegelNeurIPS 2025 · 被引用 7 次
- Beyond Labels and Topics: Discovering Causal Relationships in Neural Topic ModelingYi-Kun Tang, Heyan Huang, Xuewen Shi, Xian-Ling MaoWWW 2024 · 被引用 4 次
- Comparison of meta-learners for estimating multi-valued treatment heterogeneous effectsNaoufal Acharki, Ramiro Lugo, Antoine Bertoncello, Josselin GarnierICML 2023 · 被引用 18 次
- E-LDA: Toward Interpretable LDA Topic Models with Strong Guarantees in Logarithmic Parallel TimeAdam BreuerICML 2025
