PluRel: Synthetic Data unlocks Scaling Laws for Relational Foundation Models
Vignesh Kothapalli, Rishabh Ranjan, Valter Hudovernik, Vijay Prakash Dwivedi, Johannes Hoffart, Carlos Guestrin, Jure Leskovec
摘要
Relational Foundation Models (RFMs) facilitate data-driven decision-making by learning from complex multi-table databases. However, the diverse relational databases needed to train such models are rarely public due to privacy constraints. While there are methods to generate synthetic tabular data of arbitrary size, incorporating schema structure and primary-foreign key connectivity for multi-table generation remains challenging. Here we introduce PLUREL, a framework to synthesize multi-tabular relational databases from scratch. In a step-by-step fashion, PLUREL models (1) schemas with directed graphs, (2) intertable primary-foreign key connectivity with bipartite graphs, and, (3) feature distributions in tables via conditional causal mechanisms. The design space across these stages supports the synthesis of a wide range of diverse databases, while being computationally lightweight. Using PLUREL, we observe for the first time that (1) RFM pretraining loss exhibits power-law scaling with the number of synthetic databases and total pretraining tokens, (2) scaling the number of synthetic databases improves generalization to real databases, and (3) synthetic pretraining yields strong base models for continued pretraining on real databases. Overall, our framework and results position synthetic data scaling as a promising paradigm for RFMs. Webpage: https://star-project.stanford.edu/plurel/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Scaling Vision TransformersXiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas BeyerCVPR 2022 · 被引用 767 次
- Transformers Can Do Bayesian InferenceSamuel Müller, Noah Hollmann, Sebastian Pineda-Arango, Josif Grabocka 等ICLR 2022 · 被引用 287 次
- Small-scale proxies for large-scale Transformer training instabilitiesMitchell Wortsman, Peter J. Liu, Lechao Xiao, Katie E. Everett 等ICLR 2024 · 被引用 162 次
- ClavaDDPM: Multi-relational Data Synthesis with Cluster-guided Diffusion ModelsWei Pang, Masoumeh Shafieinejad, Lucy Liu, Stephanie Hazlewood 等NeurIPS 2024 · 被引用 39 次
- ConTextTab: A Semantics-Aware Tabular In-Context LearnerMarco Spinaci, Marek Polewczyk, Maximilian Schambach, Sam ThelinNeurIPS 2025 · 被引用 36 次
相关 Paper
- Relational In-Context Learning via Synthetic Pre-training with Structural PriorYanbo Wang, Jiaxuan You, Chuan Shi, Muhan ZhangICML 2026 · 被引用 8 次
- IRG: Modular Synthetic Relational Database Generation with Complex Relational SchemasJiayu Li, Zilong Zhao, Milad Abdollahzadeh, Biplab Sikdar 等KDD 2026
- PrivLava: Synthesizing Relational Data with Foreign Keys under Differential PrivacyKuntai Cai, Xiaokui Xiao, Graham CormodeSIGMOD 2023 · 被引用 25 次
- Graph-Conditional Flow Matching for Relational Data GenerationDavide Scassola, Sebastiano Saccani, Luca BortolussiAAAI 2026 · 被引用 3 次
- CTSyn: A Foundation Model for Cross Tabular Data GenerationXiaofeng Lin, Chenheng Xu, Matthew Yang, Guang ChengICLR 2025
