AugMask: Training Diffusion Models on Incomplete Tabular Data via Stochastic Augmentation and Masking
Jungkyu Kim, Taeyoung Park, Kibok Lee
Abstract
Score-based diffusion models have emerged as prominent deep generative models; however, their application to tabular data remains challenging because their backbones assume fully specified inputs, whereas real-world tabular data often contain missing values. We propose AugMask , a plug-and-play training framework that adapts missing-unaware backbones to incomplete data via stochastic regularization. AugMask 1) completes inputs via conditional stochastic augmentation using lightweight auxiliary models and 2) masks the loss, using augmented missing entries for conditioning while restricting supervision to observed coordinates. We connect AugMask to a Rao-Blackwellized objective and show that marginalizing missing entries yields a variance-weighted sensitivity penalty, promoting invariance of observed-coordinate reconstruction with respect to uncertain missing entries. Across diverse datasets and missingness regimes, AugMask enables standard diffusion-based tabular generators to match or outperform specialized missing-aware baselines in both sample fidelity and downstream utility. The code will be released.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f2e40c53-ee07-475f-a11e-b48ec49c4b2bBuilds on13
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 3,959 citations
- Maximum Likelihood Training of Score-Based Diffusion ModelsYang Song, Conor Durkan, Iain Murray, Stefano ErmonNeurIPS 2021 · 958 citations
- How Faithful is your Synthetic Data? Sample-level Metrics for Evaluating and Auditing Generative ModelsAhmed M. Alaa, Boris van Breugel, Evgeny S. Saveliev, Mihaela van der SchaarICML 2022 · 287 citations
- Mixed-Type Tabular Data Synthesis with Score-based Diffusion in Latent SpaceHengrui Zhang, Jiani Zhang, Zhengyuan Shen, Balasubramaniam Srinivasan et al.ICLR 2024 · 233 citations
- VAEM: a Deep Generative Model for Heterogeneous Mixed Type DataChao Ma, Sebastian Tschiatschek, Richard E. Turner, José Miguel Hernández-Lobato et al.NeurIPS 2020 · 105 citations
Related papers
- Active Tabular Augmentation via Policy-Guided Diffusion InpaintingZheyu Zhang, Shuo Yang, Bardh Prenkaj, Gjergji KasneciICML 2026
- ReMasker: Imputing Tabular Data with Masked AutoencodingTianyu Du, Luca Melis, Ting WangICLR 2024 · 41 citations
- Diffusion Transformers for Tabular Data Time Series GenerationFabrizio Garuti, Enver Sangineto, Simone Luetto, Lorenzo Forni et al.ICLR 2025 · 12 citations
- Towards Synthesizing High-Dimensional Tabular Data with Limited SamplesZuqing Li, Junhao Gan, Jianzhong QiAAAI 2026
- TabDiff: a Mixed-type Diffusion Model for Tabular Data GenerationJuntong Shi, Minkai Xu, Harper Hua, Hengrui Zhang et al.ICLR 2025
