SAGE: Sparse Adaptive Guidance for Dependency-Aware Tabular Data Generation
Shuo Yang, Zheyu Zhang, Bardh Prenkaj, Gjergji Kasneci
Abstract
Generating high-fidelity synthetic tabular data remains a critical challenge for enhancing data availability in privacy-sensitive and lowresource domains. Recent approaches leverage LLMs by representing table rows as sequences, yet suffer from two fundamental limitations: (1) they model feature dependencies densely, introducing spurious correlations; and (2) they assume static relationships between features, ignoring how these dependencies vary with feature values. To overcome these limitations, we introduce SAGE (Sparse Adaptive Guidance), a novel LLM-based generation framework that enforces sparse and dynamic dependency guidance. SAGE discretizes features into value-aware pseudo-features and constructs a mutual information-based sparse dependency graph. This graph adaptively guides generation through explicit context selection or implicit logit correction, enabling LLMs to focus on truly relevant information during synthesis. Our extensive experiments across six datasets and multiple tasks reveal that SAGE not only improves data fidelity and downstream utility, boosting F1 scores by 10% compared to previous LLM-based methods, but also reduces policy violations by one point. These results highlight the importance of adaptive structure in tabular data generation and provide new insights into context-sensitive control of LLMs. 1 * Equal contribution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3ab7863f-b1df-4a56-9b5c-2473a811a27eBuilds on7
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- TabDDPM: Modelling Tabular Data with Diffusion ModelsAkim Kotelnikov, Dmitry Baranchuk, Ivan Rubachev, Artem BabenkoICML 2023 · 518 citations
- Mixed-Type Tabular Data Synthesis with Score-based Diffusion in Latent SpaceHengrui Zhang, Jiani Zhang, Zhengyuan Shen, Balasubramaniam Srinivasan et al.ICLR 2024 · 233 citations
- Language Models are Realistic Tabular Data GeneratorsVadim Borisov, Kathrin Seßler, Tobias Leemann, Martin Pawelczyk et al.ICLR 2023 · 45 citations
- EPIC: Effective Prompting for Imbalanced-Class Data Synthesis in Tabular Data Classification via Large Language ModelsJinhee Kim, Taesung Kim, Jaegul ChooNeurIPS 2024 · 22 citations
Related papers
- Doubling Your Data in Minutes: Ultra-fast Tabular Data Generation via LLM-Induced Dependency GraphsShuo Yang, Zheyu Zhang, Bardh Prenkaj, Gjergji KasneciEMNLP 2025 · 6 citations
- LCATS: LLM-Guided Constraint-Aware Tabular Data SynthesisQing Li, Yanyan Shen, Qibin Zheng, Yi Liu et al.KDD 2026
- PrAda-GAN: A Private Adaptive Generative Adversarial Network with Bayes Network StructureKe Jia, Yuheng Ma, Yang Li, Feifei WangAAAI 2026
- ReTabSyn: Realistic Tabular Data Synthesis via Reinforcement LearningXiaofeng Lin, Seungbae Kim, Zhuoya Li, Zachary DeSoto et al.ICML 2026
- AFT-Tab: Adversarial Fine-Tuning for Tabular Data Synthesis with Long Text ColumnsYuhao Zhang, Liang Yan, Shaoming Duan, Xinyu Zha et al.ACL 2026
