CuTS: Customizable Tabular Synthetic Data Generation
Mark Vero, Mislav Balunovic, Martin T. Vechev
摘要
Privacy, data quality, and data sharing concerns pose a key limitation for tabular data applications. While generating synthetic data resembling the original distribution addresses some of these issues, most applications would benefit from additional customization on the generated data. However, existing synthetic data approaches are limited to particular constraints, e.g., differential privacy (DP) or fairness. In this work, we introduce CuTS, the first customizable synthetic tabular data generation framework. Customization in CuTS is achieved via declarative statistical and logical expressions, supporting a wide range of requirements (e.g., DP or fairness, among others). To ensure high synthetic data quality in the presence of custom specifications, CuTS is pre-trained on the original dataset and fine-tuned on a differentiable loss automatically derived from the provided specifications using novel relaxations. We evaluate CuTS over four datasets and on numerous custom specifications, outperforming state-of-the-art specialized approaches on several tasks while being more general. In particular, at the same fairness level, we achieve 2.3% higher downstream accuracy than the state-of-the-art in fair synthetic data generation on the Adult dataset.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- ReTabSyn: Realistic Tabular Data Synthesis via Reinforcement LearningXiaofeng Lin, Seungbae Kim, Zhuoya Li, Zachary DeSoto 等ICML 2026
- The Importance of Being Discrete: Measuring the Impact of Discretization in End-to-End Differentially Private Synthetic DataGeorgi Ganev, Meenatchi Sundaram Muthu Selva Annamalai, Sofiane Mahiou, Emiliano De CristofaroCCS 2025
- KGMark: A Diffusion Watermark for Knowledge GraphsHongrui Peng, Haolang Lu, Yuanlong Yu, Weiye Fu 等ICML 2025
它引用的顶会 Paper19
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan 等CCS 2016 · 被引用 7,620 次
- DECAF: Generating Fair Synthetic Data Using Causally-Aware Generative NetworksBoris van Breugel, Trent Kyono, Jeroen Berrevoets, Mihaela van der SchaarNeurIPS 2021 · 被引用 174 次
- AIM: An Adaptive and Iterative Mechanism for Differentially Private Synthetic DataRyan McKenna, Brett Mullins, Daniel Sheldon, Gerome MiklauVLDB 2022 · 被引用 136 次
- Semantic Probabilistic Layers for Neuro-Symbolic LearningKareem Ahmed, Stefano Teso, Kai-Wei Chang, Guy Van den Broeck 等NeurIPS 2022 · 被引用 133 次
- Iterative Methods for Private Synthetic Data: Unifying Framework and New MethodsTerrance Liu, Giuseppe Vietri, Steven WuNeurIPS 2021 · 被引用 85 次
相关 Paper
- PreFair: Privately Generating Justifiably Fair Synthetic DataDavid Pujol, Amir Gilad, Ashwin MachanavajjhalaVLDB 2023 · 被引用 16 次
- Differentially Private Synthetic Data via APIs 4: Tabular DataToan Tran, Arturs Backurs, Zinan Lin, Victor Reis 等ICML 2026 · 被引用 1 次
- Graphical vs. Deep Generative Models: Measuring the Impact of Differentially Private Mechanisms and Budgets on UtilityGeorgi Ganev, Kai Xu, Emiliano De CristofaroCCS 2024 · 被引用 5 次
- SynDiSC: High-Quality Tabular Data Synthesis with Distributional and Semantic ConsistencyFan Wu, Haoye Pan, Hao Wu, Kai Qian 等SIGIR 2026 · 被引用 1 次
- LCATS: LLM-Guided Constraint-Aware Tabular Data SynthesisQing Li, Yanyan Shen, Qibin Zheng, Yi Liu 等KDD 2026
