Towards Causal Foundation Model: on Duality between Optimal Balancing and Attention
Jiaqi Zhang, Joel Jennings, Agrin Hilmkil, Nick Pawlowski, Cheng Zhang, Chao Ma
Abstract
Foundation models have brought changes to the landscape of machine learning, demonstrating sparks of human-level intelligence across a diverse array of tasks. However, a gap persists in complex tasks such as causal inference, primarily due to challenges associated with intricate reasoning steps and high numerical precision requirements. In this work, we take a first step towards building causally-aware foundation models for treatment effect estimations. We propose a novel, theoretically justified method called Causal Inference with Attention (CInA), which utilizes multiple unlabeled datasets to perform self-supervised causal learning, and subsequently enables zeroshot causal inference on unseen tasks with new data. This is based on our theoretical results that demonstrate the primal-dual connection between optimal covariate balancing and self-attention, facilitating zero-shot causal inference through the final layer of a trained transformer-type architecture. We demonstrate empirically that CInA effectively generalizes to out-of-distribution datasets and various real-world datasets, matching or even surpassing traditional per-dataset methodologies. These results provide compelling evidence that our method has the potential to serve as a stepping stone for the development of causal foundation models. * Equal contribution 1 Massachusetts Institute of Technology 2 Google DeepMind 3 Work done while at Microsoft
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5a4bde2b-3487-44e3-a008-d5eccd53397eCited by top-tier papers4
- Foundation Models for Causal Inference via Prior-Data Fitted NetworksYuchen Ma, Dennis Frauen, Emil Javurek, Stefan FeuerriegelICLR 2026 · 37 citations
- Estimating Interventional Distributions with Uncertain Causal Graphs through Meta-LearningAnish Dhir, Cristiana Diaconu, Valentinian Lungu, James Requeima et al.NeurIPS 2025 · 16 citations
- Controllable Sequence Editing for Biological and Clinical TrajectoriesMichelle M. Li, Kevin Li, Yasha Ektefaie, Ying Jin et al.ICLR 2026
- A Generalization Theory for Zero-Shot PredictionRonak Mehta, Zaïd HarchaouiICML 2025
Builds on16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- A decoder-only foundation model for time-series forecastingAbhimanyu Das, Weihao Kong, Rajat Sen, Yichen ZhouICML 2024 · 601 citations
Related papers
- Causal Attention for Unbiased Visual RecognitionTan Wang, Chang Zhou, Qianru Sun, Hanwang ZhangICCV 2021 · 162 citations
- Causal Interpretation of Self-Attention in Pre-Trained TransformersRaanan Y. Rohekar, Yaniv Gurwicz, Shami NisimovNeurIPS 2023 · 62 citations
- Let Go of Your Labels with Unsupervised TransferArtyom Gadetsky, Yulun Jiang, Maria BrbicICML 2024 · 16 citations
- Use What You Know: Causal Foundation Models with Partial GraphsArik Reuter, Anish Dhir, Cristiana Diaconu, Jake Robertson et al.ICML 2026 · 2 citations
- Cross-Lingual Transfer with Class-Weighted Language-Invariant RepresentationsRuicheng Xian, Heng Ji, Han ZhaoICLR 2022 · 5 citations
