Doubling Your Data in Minutes: Ultra-fast Tabular Data Generation via LLM-Induced Dependency Graphs
Shuo Yang, Zheyu Zhang, Bardh Prenkaj, Gjergji Kasneci
摘要
Tabular data is critical across diverse domains, yet high-quality datasets remain scarce due to privacy concerns and the cost of collection. Contemporary approaches adopt large language models (LLMs) for tabular augmentation, but exhibit two major limitations: (1) dense dependency modeling among tabular features that can introduce bias, and (2) high computational overhead in sampling. To address these issues, we propose SPADA for SPArse Dependency-driven Augmentation, a lightweight generative framework that explicitly captures sparse dependencies via an LLM-induced graph. We treat each feature as a node and synthesize values by traversing the graph, conditioning each feature solely on its parent nodes. We explore two synthesis strategies: a non-parametric method using Gaussian kernel density estimation, and a conditional normalizing flow model that learns invertible mappings for conditional density estimation. Experiments on four datasets show that SPADA reduces constraint violations by 4% compared to diffusion-based methods and accelerates generation by nearly 9,500 times over LLM-based baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Active Tabular Augmentation via Policy-Guided Diffusion InpaintingZheyu Zhang, Shuo Yang, Bardh Prenkaj, Gjergji KasneciICML 2026
- SAGE: Sparse Adaptive Guidance for Dependency-Aware Tabular Data GenerationShuo Yang, Zheyu Zhang, Bardh Prenkaj, Gjergji KasneciACL 2026
它引用的顶会 Paper6
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- TabDDPM: Modelling Tabular Data with Diffusion ModelsAkim Kotelnikov, Dmitry Baranchuk, Ivan Rubachev, Artem BabenkoICML 2023 · 被引用 518 次
- Mixed-Type Tabular Data Synthesis with Score-based Diffusion in Latent SpaceHengrui Zhang, Jiani Zhang, Zhengyuan Shen, Balasubramaniam Srinivasan 等ICLR 2024 · 被引用 233 次
- Language Models are Realistic Tabular Data GeneratorsVadim Borisov, Kathrin Seßler, Tobias Leemann, Martin Pawelczyk 等ICLR 2023 · 被引用 45 次
- Beyond the convexity assumption: Realistic tabular data generation under quantifier-free real linear constraintsMihaela C. Stoian, Eleonora GiunchigliaICLR 2025
相关 Paper
- LCATS: LLM-Guided Constraint-Aware Tabular Data SynthesisQing Li, Yanyan Shen, Qibin Zheng, Yi Liu 等KDD 2026
- TabLoft: Tabular Data Generation Based on LLM with Ordered FeaturesLuyu Chen, Changhao Wu, Jingyi Li, Sen Liu 等ICDE 2026
- Counterfactual Residual Data Augmentation for RegressionHossein Mohebbi, Oliver Schulte, Ke Li, Pascal PoupartICML 2026
- PrAda-GAN: A Private Adaptive Generative Adversarial Network with Bayes Network StructureKe Jia, Yuheng Ma, Yang Li, Feifei WangAAAI 2026
- Latent Diffusion-based Data Augmentation for Continuous-Time Dynamic Graph ModelYuxing Tian, Aiwen Jiang, Qi Huang, Jian Guo 等KDD 2024 · 被引用 7 次
