Use What You Know: Causal Foundation Models with Partial Graphs
Arik Reuter, Anish Dhir, Cristiana Diaconu, Jake Robertson, Ole Ossen, Frank Hutter, Adrian Weller, Mark van der Wilk, Bernhard Schölkopf
Abstract
Estimating causal quantities traditionally relies on bespoke estimators tailored to specific assumptions. Recently proposed Causal Foundation Models (CFMs) promise a more unified approach by amortising causal discovery and inference in a single step. However, in their current state, they do not allow for the incorporation of any domain knowledge, which can lead to suboptimal predictions. We bridge this gap by introducing methods to condition CFMs on causal information, such as the causal graph or more readily available ancestral information. When access to complete causal graph information is too strict a requirement, our approach also effectively leverages partial causal information. We systematically evaluate conditioning strategies and find that injecting learnable biases into the attention mechanism, together with a graph-convolutional encoder, is a highly effective method to utilise full and partial causal information. Our experiments show that this conditioning allows a general-purpose CFM to match the performance of specialised models trained on specific causal structures. Overall, our approach addresses a central hurdle on the path towards allin-one causal foundation models: the capability to answer causal queries in a data-driven manner while effectively leveraging any amount of domain expertise.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bba2cde6-a8c5-4121-94b2-616d885f4a34Builds on15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Transformers Can Do Bayesian InferenceSamuel Müller, Noah Hollmann, Sebastian Pineda-Arango, Josif Grabocka et al.ICLR 2022 · 287 citations
- Transformer Neural Processes: Uncertainty-Aware Meta Learning Via Sequence ModelingTung Nguyen, Aditya GroverICML 2022 · 148 citations
Related papers
- Foundation Models for Causal Inference via Prior-Data Fitted NetworksYuchen Ma, Dennis Frauen, Emil Javurek, Stefan FeuerriegelICLR 2026 · 37 citations
- Amortized Inference for Causal Structure LearningLars Lorch, Scott Sussex, Jonas Rothfuss, Andreas Krause et al.NeurIPS 2022 · 118 citations
- Conditional Independence Testing with Heteroskedastic Data and Applications to Causal DiscoveryWiebke Günther, Urmi Ninad, Jonas Wahl, Jakob RungeNeurIPS 2022 · 6 citations
- Towards Causal Foundation Model: on Duality between Optimal Balancing and AttentionJiaqi Zhang, Joel Jennings, Agrin Hilmkil, Nick Pawlowski et al.ICML 2024 · 9 citations
- Counterfactual Fairness with Partially Known Causal GraphAoqi Zuo, Susan Wei, Tongliang Liu, Bo Han et al.NeurIPS 2022 · 32 citations
