STaSy: Score-based Tabular data Synthesis
Jayoung Kim, Chaejeong Lee, Noseong Park
摘要
Tabular data synthesis is a long-standing research topic in machine learning. Many different methods have been proposed over the past decades, ranging from statistical methods to deep generative methods. However, it has not always been successful due to the complicated nature of real-world tabular data. In this paper, we present a new model named Score-based Tabular data Synthesis (STaSy) and its training strategy based on the paradigm of score-based generative modeling. Despite the fact that score-based generative models have resolved many issues in generative models, there still exists room for improvement in tabular data synthesis. Our proposed training strategy includes a self-paced learning technique and a fine-tuning strategy, which further increases the sampling quality and diversity by stabilizing the denoising score matching training. Furthermore, we also conduct rigorous experimental studies in terms of the generative task trilemma: sampling quality, diversity, and time. In our experiments with 15 benchmark tabular datasets and 7 baselines, our method outperforms existing methods in terms of task-dependant evaluations and diversity. Code is available at https://github.com/JayoungKim408/STaSy .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- Mixed-Type Tabular Data Synthesis with Score-based Diffusion in Latent SpaceHengrui Zhang, Jiani Zhang, Zhengyuan Shen, Balasubramaniam Srinivasan 等ICLR 2024 · 被引用 233 次
- Curated LLM: Synergy of LLMs and Data Curation for tabular augmentation in low-data regimesNabeel Seedat, Nicolas Huynh, Boris van Breugel, Mihaela van der SchaarICML 2024 · 被引用 61 次
- ClavaDDPM: Multi-relational Data Synthesis with Cluster-guided Diffusion ModelsWei Pang, Masoumeh Shafieinejad, Lucy Liu, Stephanie Hazlewood 等NeurIPS 2024 · 被引用 39 次
- How Realistic Is Your Synthetic Data? Constraining Deep Generative Models for Tabular DataMihaela C. Stoian, Salijona Dyrmishi, Maxime Cordy, Thomas Lukasiewicz 等ICLR 2024 · 被引用 33 次
- TabEBM: A Tabular Data Augmentation Method with Distinct Class-Specific Energy-Based ModelsAndrei Margeloiu, Xiangjian Jiang, Nikola Simidjievski, Mateja JamnikNeurIPS 2024 · 被引用 19 次
它引用的顶会 Paper5
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar 等ICLR 2021 · 被引用 1,270 次
- Tackling the Generative Learning Trilemma with Denoising Diffusion GANsZhisheng Xiao, Karsten Kreis, Arash VahdatICLR 2022 · 被引用 726 次
- Reliable Fidelity and Diversity Metrics for Generative ModelsMuhammad Ferjad Naeem, Seong Joon Oh, Youngjung Uh, Yunjey Choi 等ICML 2020 · 被引用 553 次
- How to Train Your Neural ODE: the World of Jacobian and Kinetic RegularizationChris Finlay, Jörn-Henrik Jacobsen, Levon Nurbekyan, Adam M. ObermanICML 2020 · 被引用 76 次
- Invertible Tabular GANs: Killing Two Birds with One Stone for Tabular Data SynthesisJaehoon Lee, Jihyeon Hyeong, Jinsung Jeon, Noseong Park 等NeurIPS 2021 · 被引用 39 次
相关 Paper
- SOS: Score-based Oversampling for Tabular DataJayoung Kim, Chaejeong Lee, Yehjin Shin, Sewon Park 等KDD 2022 · 被引用 20 次
- CTSyn: A Foundation Model for Cross Tabular Data GenerationXiaofeng Lin, Chenheng Xu, Matthew Yang, Guang ChengICLR 2025
- CoDi: Co-evolving Contrastive Diffusion Models for Mixed-type Tabular SynthesisChaejeong Lee, Jayoung Kim, Noseong ParkICML 2023 · 被引用 97 次
- STUNT: Few-shot Tabular Learning with Self-generated Tasks from Unlabeled TablesJaehyun Nam, Jihoon Tack, Kyungmin Lee, Hankook Lee 等ICLR 2023 · 被引用 2 次
- SynDiSC: High-Quality Tabular Data Synthesis with Distributional and Semantic ConsistencyFan Wu, Haoye Pan, Hao Wu, Kai Qian 等SIGIR 2026 · 被引用 1 次
