Pairwise Causality Guided Transformers for Event Sequences
Xiao Shou, Debarun Bhattacharjya, Tian Gao, Dharmashankar Subramanian, Oktie Hassanzadeh, Kristin P. Bennett
Abstract
Although pairwise causal relations have been extensively studied in observational longitudinal analyses across many disciplines, incorporating knowledge of causal pairs into deep learning models for temporal event sequences remains largely unexplored. In this paper, we propose a novel approach for enhancing the performance of transformer-based models in multivariate event sequences by injecting pairwise qualitative causal knowledge such as ‘event Z amplifies future occurrences of event Y’. We establish a new framework for causal inference in temporal event sequences using a transformer architecture, providing a theoretical justification for our approach, and show how to obtain unbiased estimates of the proposed measure. Experimental results demonstrate that our approach outperforms several state-of-the-art models in terms of prediction accuracy by effectively leveraging knowledge about causal pairs. We also consider a unique application where we extract knowledge around sequences of societal events by generating them from a large language model, and demonstrate how a causal knowledge graph can help with event prediction in such sequences. Overall, our framework offers a practical means of improving the performance of transformer-based models in multivariate event sequences by explicitly exploiting pairwise causal information.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 220248bb-453f-456d-b940-2463ab5a2fc6Builds on13
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Estimating counterfactual treatment outcomes over time through adversarially balanced representationsIoana Bica, Ahmed M. Alaa, James Jordon, Mihaela van der SchaarICLR 2020 · 224 citations
- Causal Transformer for Estimating Counterfactual OutcomesValentyn Melnychuk, Dennis Frauen, Stefan FeuerriegelICML 2022 · 146 citations
- CAUSE: Learning Granger Causality from Event Sequences using Attribution MethodsWei Zhang, Thomas Kobber Panum, Somesh Jha, Prasad Chalasani et al.ICML 2020 · 64 citations
Related papers
- Transformer Hawkes ProcessSimiao Zuo, Haoming Jiang, Zichong Li, Tuo Zhao et al.ICML 2020 · 382 citations
- GraFT: Infusing Pre-trained Transformers with Relational Structure for Time Series ForecastingYuqi Yuan, Xiong Luo, Qiaojuan Peng, Wenbing ZhaoAAAI 2026
- Causal Graph based Event Reasoning using Semantic Relation ExpertsMahnaz Koupaee, Xueying Bai, Mudan Chen, Greg Durrett et al.ACL 2025
- Causal Interpretation of Self-Attention in Pre-Trained TransformersRaanan Y. Rohekar, Yaniv Gurwicz, Shami NisimovNeurIPS 2023 · 62 citations
- Robust Event Forecasting with Spatiotemporal Confounder LearningSonggaojun Deng, Huzefa Rangwala, Yue NingKDD 2022 · 9 citations
