S i 1 o F use: Cross-silo Synthetic Data Generation with Latent Tabular Diffusion Models
Aditya Shankar, Hans Brouwer, Rihan Hai, Lydia Y. Chen
Abstract
Synthetic tabular data is crucial for sharing and augmenting data across silos, especially for enterprises with proprietary data. However, existing synthesizers are designed for centrally stored data. Hence, they struggle with real-world scenarios where features are distributed across multiple silos, necessitating on-premise data storage. We introduce SiloFuse, a novel generative framework for high-quality synthesis from cross-silo tabular data. To ensure privacy, SiloFuse utilizes a distributed latent tabular diffusion architecture. Through autoencoders, latent representations are learned for each client's features, masking their actual values. We employ stacked dis-tributed training to improve communication efficiency, reducing the number of rounds to a single step. Under SiloFuse, we prove the impossibility of data reconstruction for vertically partitioned synthesis and quantify privacy risks through three attacks using our benchmark framework. Experimental results on nine datasets showcase SiloFuse's competence against centralized diffusion-based synthesizers. Notably, SiloFuse achieves 43.8 and 29.8 higher percentage points over GANs in resemblance and utility. Experiments on communication show stacked training's fixed cost compared to the growing costs of end-to-end training as the number of training iterations increases. Additionally, SiloFuse proves robust to feature permutations and varying numbers of clients.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- WaveStitch: Flexible and Fast Conditional Time Series Generation With Diffusion ModelsAditya Shankar, Lydia Yiyu Chen, Arie van Deursen, Rihan HaiSIGMOD 2026 · 3 citations
- Harpoon: Generalised Manifold Guidance for Conditional Tabular DiffusionAditya Shankar, Yuandou Wang, Rihan Hai, Lydia Y. ChenICLR 2026 · 2 citations
Builds on9
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Inverting Gradients - How easy is it to break privacy in federated learning?Jonas Geiping, Hartmut Bauermeister, Hannah Dröge, Michael MoellerNeurIPS 2020 · 1,822 citations
- Argmax Flows and Multinomial Diffusion: Learning Categorical DistributionsEmiel Hoogeboom, Didrik Nielsen, Priyank Jaini, Patrick Forré et al.NeurIPS 2021 · 782 citations
Related papers
- CTSyn: A Foundation Model for Cross Tabular Data GenerationXiaofeng Lin, Chenheng Xu, Matthew Yang, Guang ChengICLR 2025
- Systematic Assessment of Tabular Data SynthesisYuntao Du, Ninghui LiCCS 2025 · 2 citations
- Invertible Tabular GANs: Killing Two Birds with One Stone for Tabular Data SynthesisJaehoon Lee, Jihyeon Hyeong, Jinsung Jeon, Noseong Park et al.NeurIPS 2021 · 39 citations
- FLAIM: AIM-based Synthetic Data Generation in the Federated SettingSamuel Maddock, Graham Cormode, Carsten MapleKDD 2024 · 5 citations
- HeteroFedSyn: Differentially Private Tabular Data Synthesis for Heterogeneous Federated SettingsXiaochen Li, Fengyu Gao, Xizixiang Wei, Tianhao Wang et al.SIGMOD 2026
