CuTS: Customizable Tabular Synthetic Data Generation
Mark Vero, Mislav Balunovic, Martin T. Vechev
Abstract
Privacy, data quality, and data sharing concerns pose a key limitation for tabular data applications. While generating synthetic data resembling the original distribution addresses some of these issues, most applications would benefit from additional customization on the generated data. However, existing synthetic data approaches are limited to particular constraints, e.g., differential privacy (DP) or fairness. In this work, we introduce CuTS, the first customizable synthetic tabular data generation framework. Customization in CuTS is achieved via declarative statistical and logical expressions, supporting a wide range of requirements (e.g., DP or fairness, among others). To ensure high synthetic data quality in the presence of custom specifications, CuTS is pre-trained on the original dataset and fine-tuned on a differentiable loss automatically derived from the provided specifications using novel relaxations. We evaluate CuTS over four datasets and on numerous custom specifications, outperforming state-of-the-art specialized approaches on several tasks while being more general. In particular, at the same fairness level, we achieve 2.3% higher downstream accuracy than the state-of-the-art in fair synthetic data generation on the Adult dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f8294d7f-5d20-46da-bec6-7681a6905359Cited by top-tier papers3
- ReTabSyn: Realistic Tabular Data Synthesis via Reinforcement LearningXiaofeng Lin, Seungbae Kim, Zhuoya Li, Zachary DeSoto et al.ICML 2026
- The Importance of Being Discrete: Measuring the Impact of Discretization in End-to-End Differentially Private Synthetic DataGeorgi Ganev, Meenatchi Sundaram Muthu Selva Annamalai, Sofiane Mahiou, Emiliano De CristofaroCCS 2025
- KGMark: A Diffusion Watermark for Knowledge GraphsHongrui Peng, Haolang Lu, Yuanlong Yu, Weiye Fu et al.ICML 2025
Builds on19
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- DECAF: Generating Fair Synthetic Data Using Causally-Aware Generative NetworksBoris van Breugel, Trent Kyono, Jeroen Berrevoets, Mihaela van der SchaarNeurIPS 2021 · 174 citations
- AIM: An Adaptive and Iterative Mechanism for Differentially Private Synthetic DataRyan McKenna, Brett Mullins, Daniel Sheldon, Gerome MiklauVLDB 2022 · 136 citations
- Semantic Probabilistic Layers for Neuro-Symbolic LearningKareem Ahmed, Stefano Teso, Kai-Wei Chang, Guy Van den Broeck et al.NeurIPS 2022 · 133 citations
- Iterative Methods for Private Synthetic Data: Unifying Framework and New MethodsTerrance Liu, Giuseppe Vietri, Steven WuNeurIPS 2021 · 85 citations
Related papers
- PreFair: Privately Generating Justifiably Fair Synthetic DataDavid Pujol, Amir Gilad, Ashwin MachanavajjhalaVLDB 2023 · 16 citations
- Differentially Private Synthetic Data via APIs 4: Tabular DataToan Tran, Arturs Backurs, Zinan Lin, Victor Reis et al.ICML 2026 · 1 citation
- Graphical vs. Deep Generative Models: Measuring the Impact of Differentially Private Mechanisms and Budgets on UtilityGeorgi Ganev, Kai Xu, Emiliano De CristofaroCCS 2024 · 5 citations
- SynDiSC: High-Quality Tabular Data Synthesis with Distributional and Semantic ConsistencyFan Wu, Haoye Pan, Hao Wu, Kai Qian et al.SIGIR 2026 · 1 citation
- LCATS: LLM-Guided Constraint-Aware Tabular Data SynthesisQing Li, Yanyan Shen, Qibin Zheng, Yi Liu et al.KDD 2026
