S i 1 o F use: Cross-silo Synthetic Data Generation with Latent Tabular Diffusion Models
Aditya Shankar, Hans Brouwer, Rihan Hai, Lydia Y. Chen
摘要
Synthetic tabular data is crucial for sharing and augmenting data across silos, especially for enterprises with proprietary data. However, existing synthesizers are designed for centrally stored data. Hence, they struggle with real-world scenarios where features are distributed across multiple silos, necessitating on-premise data storage. We introduce SiloFuse, a novel generative framework for high-quality synthesis from cross-silo tabular data. To ensure privacy, SiloFuse utilizes a distributed latent tabular diffusion architecture. Through autoencoders, latent representations are learned for each client's features, masking their actual values. We employ stacked dis-tributed training to improve communication efficiency, reducing the number of rounds to a single step. Under SiloFuse, we prove the impossibility of data reconstruction for vertically partitioned synthesis and quantify privacy risks through three attacks using our benchmark framework. Experimental results on nine datasets showcase SiloFuse's competence against centralized diffusion-based synthesizers. Notably, SiloFuse achieves 43.8 and 29.8 higher percentage points over GANs in resemblance and utility. Experiments on communication show stacked training's fixed cost compared to the growing costs of end-to-end training as the number of training iterations increases. Additionally, SiloFuse proves robust to feature permutations and varying numbers of clients.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- WaveStitch: Flexible and Fast Conditional Time Series Generation With Diffusion ModelsAditya Shankar, Lydia Yiyu Chen, Arie van Deursen, Rihan HaiSIGMOD 2026 · 被引用 3 次
- Harpoon: Generalised Manifold Guidance for Conditional Tabular DiffusionAditya Shankar, Yuandou Wang, Rihan Hai, Lydia Y. ChenICLR 2026 · 被引用 2 次
它引用的顶会 Paper9
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Inverting Gradients - How easy is it to break privacy in federated learning?Jonas Geiping, Hartmut Bauermeister, Hannah Dröge, Michael MoellerNeurIPS 2020 · 被引用 1,822 次
- Argmax Flows and Multinomial Diffusion: Learning Categorical DistributionsEmiel Hoogeboom, Didrik Nielsen, Priyank Jaini, Patrick Forré 等NeurIPS 2021 · 被引用 782 次
相关 Paper
- CTSyn: A Foundation Model for Cross Tabular Data GenerationXiaofeng Lin, Chenheng Xu, Matthew Yang, Guang ChengICLR 2025
- Systematic Assessment of Tabular Data SynthesisYuntao Du, Ninghui LiCCS 2025 · 被引用 2 次
- Invertible Tabular GANs: Killing Two Birds with One Stone for Tabular Data SynthesisJaehoon Lee, Jihyeon Hyeong, Jinsung Jeon, Noseong Park 等NeurIPS 2021 · 被引用 39 次
- FLAIM: AIM-based Synthetic Data Generation in the Federated SettingSamuel Maddock, Graham Cormode, Carsten MapleKDD 2024 · 被引用 5 次
- HeteroFedSyn: Differentially Private Tabular Data Synthesis for Heterogeneous Federated SettingsXiaochen Li, Fengyu Gao, Xizixiang Wei, Tianhao Wang 等SIGMOD 2026
