A Pre-training Framework for Relational Data with Information-theoretic Principles
Quang Truong, Zhikai Chen, Mingxuan Ju, Tong Zhao, Neil Shah, Jiliang Tang
Abstract
Relational databases underpin critical infrastructure across a wide range of domains, yet the design of generalizable pre-training strategies for learning from relational databases remains an open challenge due to task heterogeneity. Specifically, there exist many possible downstream tasks, as tasks are defined based on relational schema graphs, temporal dependencies, and SQL-defined label logics. An effective pre-training framework is desired to take these factors into account in order to obtain task-aware representations. By incorporating knowledge of the underlying distribution that drives label generation, downstream tasks can benefit from relevant side-channel information. To bridge this gap, we introduce Task Vector Estimation (TVE), a novel pre-training framework that constructs predictive supervisory signals via set-based aggregation over schema traversal graphs, explicitly modeling next-window relational dynamics. We formalize our approach through an information-theoretic lens, demonstrating that task-informed representations retain more relevant signals than those obtained without task priors. Extensive experiments on the RelBench benchmark show that TVE consistently outperforms traditional pre-training baselines. Our findings advocate for pre-training objectives that encode task heterogeneity and temporal structure as design principles for predictive modeling on relational databases. Our code is publicly available at https://github.com/quang-truong/task-vector-estimation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d9c57ef9-198f-40d1-b1ed-cc908808ac3dCited by top-tier papers2
- Is Fixing Schema Graphs Necessary? Full-Resolution Graph Structure Learning for Relational Deep LearningYi Huang, Qingyun Sun, Jia Li, Xingcheng Fu et al.ICML 2026
- Ramba: Selective State-Space Models for Relational Deep LearningYiming Liu, Chunyu Wei, haozhe lin, Fengjun Xiao et al.ICML 2026
Builds on16
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Graph Contrastive Learning with AugmentationsYuning You, Tianlong Chen, Yongduo Sui, Ting Chen et al.NeurIPS 2020 · 3,042 citations
- Revisiting Deep Learning Models for Tabular DataYury Gorishniy, Ivan Rubachev, Valentin Khrulkov, Artem BabenkoNeurIPS 2021 · 1,847 citations
- GraphMAE: Self-Supervised Masked Graph AutoencodersZhenyu Hou, Xiao Liu, Yukuo Cen, Yuxiao Dong et al.KDD 2022 · 533 citations
- Task2Vec: Task Embedding for Meta-LearningAlessandro Achille, Michael Lam, Rahul Tewari, Avinash Ravichandran et al.ICCV 2019 · 359 citations
Related papers
- Relational Transformer: Toward Zero-Shot Foundation Models for Relational DataRishabh Ranjan, Valter Hudovernik, Mark Znidar, Charilaos I. Kanatsoulis et al.ICLR 2026 · 35 citations
- TURL: Table Understanding through Representation LearningXiang Deng, Huan Sun, Alyssa Lees, You Wu et al.VLDB 2021 · 2,406 citations
- MUG: Meta-path-aware Universal Heterogeneous Graph Pre-TrainingLianze Shan, Jitao Zhao, Dongxiao He, Yongqi Huang et al.AAAI 2026 · 1 citation
- Relational In-Context Learning via Synthetic Pre-training with Structural PriorYanbo Wang, Jiaxuan You, Chuan Shi, Muhan ZhangICML 2026 · 8 citations
- Griffin: Towards a Graph-Centric Relational Database Foundation ModelYanbo Wang, Xiyuan Wang, Quan Gan, Minjie Wang et al.ICML 2025
