Mixed-Type Tabular Data Synthesis with Score-based Diffusion in Latent Space
Hengrui Zhang, Jiani Zhang, Zhengyuan Shen, Balasubramaniam Srinivasan, Xiao Qin, Christos Faloutsos, Huzefa Rangwala, George Karypis
摘要
Recent advances in tabular data generation have greatly enhanced synthetic data quality. However, extending diffusion models to tabular data is challenging due to the intricately varied distributions and a blend of data types of tabular data. This paper introduces TABSYN, a methodology that synthesizes tabular data by leveraging a diffusion model within a variational autoencoder (VAE) crafted latent space. The key advantages of the proposed TABSYN include (1) Generality: the ability to handle a broad spectrum of data types by converting them into a single unified space and explicitly capturing inter-column relations; (2) Quality: optimizing the distribution of latent embeddings to enhance the subsequent training of diffusion models, which helps generate high-quality synthetic data, (3) Speed: much fewer number of reverse steps and faster synthesis speed than existing diffusion-based methods. Extensive experiments on six datasets with five metrics demonstrate that TABSYN outperforms existing methods. Specifically, it reduces the error rates by 86% and 67% for column-wise distribution and pair-wise column correlation estimations compared with the most competitive baselines. The code has been made available at https://github.com/amazon-science/tabsyn .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper54
- ClavaDDPM: Multi-relational Data Synthesis with Cluster-guided Diffusion ModelsWei Pang, Masoumeh Shafieinejad, Lucy Liu, Stephanie Hazlewood 等NeurIPS 2024 · 被引用 39 次
- TabEBM: A Tabular Data Augmentation Method with Distinct Class-Specific Energy-Based ModelsAndrei Margeloiu, Xiangjian Jiang, Nikola Simidjievski, Mateja JamnikNeurIPS 2024 · 被引用 19 次
- Diffusion Transformers for Tabular Data Time Series GenerationFabrizio Garuti, Enver Sangineto, Simone Luetto, Lorenzo Forni 等ICLR 2025 · 被引用 12 次
- Joint Relational Database Generation via Graph-Conditional Diffusion ModelsMohamed Amine Ketata, David Lüdke, Leo Schwinn, Stephan GünnemannNeurIPS 2025 · 被引用 11 次
- TabStruct: Measuring Structural Fidelity of Tabular DataXiangjian Jiang, Nikola Simidjievski, Mateja JamnikICLR 2026 · 被引用 10 次
它引用的顶会 Paper21
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 被引用 3,959 次
- Revisiting Deep Learning Models for Tabular DataYury Gorishniy, Ivan Rubachev, Valentin Khrulkov, Artem BabenkoNeurIPS 2021 · 被引用 1,847 次
相关 Paper
- TabDiff: a Mixed-type Diffusion Model for Tabular Data GenerationJuntong Shi, Minkai Xu, Harper Hua, Hengrui Zhang 等ICLR 2025
- CTSyn: A Foundation Model for Cross Tabular Data GenerationXiaofeng Lin, Chenheng Xu, Matthew Yang, Guang ChengICLR 2025
- Language Models are Realistic Tabular Data GeneratorsVadim Borisov, Kathrin Seßler, Tobias Leemann, Martin Pawelczyk 等ICLR 2023 · 被引用 45 次
- SynDiSC: High-Quality Tabular Data Synthesis with Distributional and Semantic ConsistencyFan Wu, Haoye Pan, Hao Wu, Kai Qian 等SIGIR 2026 · 被引用 1 次
- TabDDPM: Modelling Tabular Data with Diffusion ModelsAkim Kotelnikov, Dmitry Baranchuk, Ivan Rubachev, Artem BabenkoICML 2023 · 被引用 518 次
