Relatron: Automating Relational Machine Learning over Relational Databases
Zhikai Chen, Han Xie, Jian Zhang, Jiliang Tang, Xiang song, Huzefa Rangwala
Abstract
Predictive modeling over relational databases (RDBs) powers applications in various domains, yet remains challenging due to the need to capture both cross-table dependencies and complex feature interactions. Recent Relational Deep Learning (RDL) methods automate feature engineering via message passing, while classical approaches like Deep Feature Synthesis (DFS) rely on predefined non-parametric aggregators. Despite promising performance gains, the comparative advantages of RDL over DFS and the design principles for selecting effective architectures remain poorly understood. We present a comprehensive study that unifies RDL and DFS in a shared design space and conducts large-scale architecture-centric searches across diverse RDB tasks. Our analysis yields three key findings: (1) RDL does not consistently outperform DFS, with performance being highly task-dependent; (2) no single architecture dominates across tasks, underscoring the need for task-aware model selection; and (3) validation accuracy is an unreliable guide for architecture choice. This search yields a curated model performance bank that links model architecture configurations to their performance; leveraging this bank, we analyze the drivers of the RDL–DFS performance gap and introduce two task signals—RDB task homophily and an affinity embedding that captures size, path, feature, and temporal structure—whose correlation with the gap enables principled routing. Guided by these signals, we propose Relatron, a task embedding-based meta-selector that first chooses between RDL and DFS and then prunes the within-family search to deliver strong performance. Lightweight loss-landscape metrics further guard against brittle checkpoints by preferring flatter optima. In experiments, Relatron resolves the “more tuning, worse performance” effect and, in joint hyperparameter–architecture optimization, achieves up to 18.5% improvement over strong baselines with lower computational cost than Fisher information–based alternatives. Our code is available at https://github.com/amazon-science/Automating-Relational-Machine-Learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ec60d43f-bf44-4405-9876-10cee23ebde7Builds on26
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong et al.NeurIPS 2020 · 3,935 citations
- Principal Neighbourhood Aggregation for Graph NetsGabriele Corso, Luca Cavalleri, Dominique Beaini, Pietro Liò et al.NeurIPS 2020 · 914 citations
- Neural Bellman-Ford Networks: A General Graph Neural Network Framework for Link PredictionZhaocheng Zhu, Zuobai Zhang, Louis-Pascal A. C. Xhonneux, Jian TangNeurIPS 2021 · 546 citations
- Design Space for Graph Neural NetworksJiaxuan You, Zhitao Ying, Jure LeskovecNeurIPS 2020 · 409 citations
- Task2Vec: Task Embedding for Meta-LearningAlessandro Achille, Michael Lam, Rahul Tewari, Avinash Ravichandran et al.ICCV 2019 · 359 citations
Related papers
- RelGNN: Composite Message Passing for Relational Deep LearningTianlang Chen, Charilaos I. Kanatsoulis, Jure LeskovecICML 2025
- What Makes a Desired Graph for Relational Deep Learning?Yao Cheng, Siqiang LuoICML 2026
- Relational Transformer: Toward Zero-Shot Foundation Models for Relational DataRishabh Ranjan, Valter Hudovernik, Mark Znidar, Charilaos I. Kanatsoulis et al.ICLR 2026 · 35 citations
- Is Fixing Schema Graphs Necessary? Full-Resolution Graph Structure Learning for Relational Deep LearningYi Huang, Qingyun Sun, Jia Li, Xingcheng Fu et al.ICML 2026
- Relational Graph TransformerVijay Prakash Dwivedi, Sri Jaladi, Yangyi Shen, Federico Lopez et al.ICLR 2026 · 35 citations
