KG-TREAT: Pre-training for Treatment Effect Estimation by Synergizing Patient Data with Knowledge Graphs
Ruoqi Liu, Lingfei Wu, Ping Zhang
Abstract
Treatment effect estimation (TEE) is the task of determining the impact of various treatments on patient outcomes. Current TEE methods fall short due to reliance on limited labeled data and challenges posed by sparse and high-dimensional observational patient data. To address the challenges, we introduce a novel pre-training and fine-tuning framework, KG-TREAT, which synergizes large-scale observational patient data with biomedical knowledge graphs (KGs) to enhance TEE. Unlike previous approaches, KG-TREAT constructs dual-focus KGs and integrates a deep bi-level attention synergy method for in-depth information fusion, enabling distinct encoding of treatment-covariate and outcome-covariate relationships. KG-TREAT also incorporates two pre-training tasks to ensure a thorough grounding and contextualization of patient data and KGs. Evaluation on four downstream TEE tasks shows KG-TREAT's superiority over existing methods, with an average improvement of 7% in Area under the ROC Curve (AUC) and 9% in Influence Function-based Precision of Estimating Heterogeneous Effects (IF-PEHE). The effectiveness of our estimated treatment effects is further affirmed by alignment with established randomized clinical trial findings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4d4e2f26-d1dd-453e-886a-eabfc8010784Cited by top-tier papers2
- LLaVA Needs More Knowledge: Retrieval Augmented Natural Language Generation with Knowledge Graph for Explaining Thoracic PathologiesAmeer Hamza, Abdullah, Yong Hyun Ahn, Sungyoung Lee et al.AAAI 2025 · 14 citations
- HLMEA: Unsupervised Entity Alignment Based on Hybrid Language ModelsXiongnan Jin, Zhilin Wang, Jinpeng Chen, Liu Yang et al.AAAI 2025 · 5 citations
Builds on8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LinkBERT: Pretraining Language Models with Document LinksMichihiro Yasunaga, Jure Leskovec, Percy LiangACL 2022 · 463 citations
- Scalable Multi-Hop Relational Reasoning for Knowledge-Aware Question AnsweringYanlin Feng, Xinyue Chen, Bill Yuchen Lin, Peifeng Wang et al.EMNLP 2020 · 207 citations
- Learning Disentangled Representations for CounterFactual RegressionNegar Hassanpour, Russell GreinerICLR 2020 · 176 citations
- On Inductive Biases for Heterogeneous Treatment Effect EstimationAlicia Curth, Mihaela van der SchaarNeurIPS 2021 · 114 citations
Related papers
- Time-aware Entity Alignment using Temporal Relational AttentionChengjin Xu, Fenglong Su, Bo Xiong, Jens LehmannWWW 2022 · 47 citations
- OntoProtein: Protein Pretraining With Gene Ontology EmbeddingNingyu Zhang, Zhen Bi, Xiaozhuan Liang, Siyuan Cheng et al.ICLR 2022 · 128 citations
- A Two-Stage Pretraining-Finetuning Framework for Treatment Effect Estimation with Unmeasured ConfoundingChuan Zhou, Yaxuan Li, Chunyuan Zheng, Haiteng Zhang et al.KDD 2025 · 6 citations
- Unsupervised Entity Alignment for Temporal Knowledge GraphsXiaoze Liu, Junyang Wu, Tianyi Li, Lu Chen et al.WWW 2023 · 56 citations
- How Knowledge Graph and Attention Help? A Qualitative Analysis into Bag-level Relation ExtractionZikun Hu, Yixin Cao, Lifu Huang, Tat-Seng ChuaACL 2021
