Masked Language Modeling Becomes Conditional Density Estimation for Tabular Data Synthesis
Seunghwan An, Gyeongdong Woo, Jaesung Lim, Chang-Hyun Kim, Sungchul Hong, Jong-June Jeon
Abstract
In this paper, our goal is to generate synthetic data for heterogeneous (mixed-type) tabular datasets with high machine learning utility (MLu). Since the MLu performance depends on accurately approximating the conditional distributions, we focus on devising a synthetic data generation method based on conditional distribution estimation. We introduce MaCoDE by redefining the consecutive multi-class classification task of Masked Language Modeling (MLM) as histogram-based non-parametric conditional density estimation. Our approach enables the estimation of conditional densities across arbitrary combinations of target and conditional variables. We bridge the theoretical gap between distributional learning and MLM by demonstrating that minimizing the orderless multi-class classification loss leads to minimizing the total variation distance between conditional distributions. To validate our proposed model, we evaluate its performance in synthetic data generation across 10 real-world datasets, demonstrating its ability to adjust data privacy levels easily without re-training. Additionally, since masked input tokens in MLM are analogous to missing data, we further assess its effectiveness in handling training datasets with missing values, including multiple imputations of the missing entries.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2a131c3d-261e-4a7d-909b-bfca766ad09aCited by top-tier papers3
- Impute Missing Entries with UncertaintyJaesung Lim, Seunghwan An, Jong-June JeonAAAI 2026
- Towards Synthesizing High-Dimensional Tabular Data with Limited SamplesZuqing Li, Junhao Gan, Jianzhong QiAAAI 2026
- TabularBERT: Binning-Based Self-Supervised Learning for Tabular RepresentationBeomjin Park, Seunghwan An, Sungchul Hong, Hosik ChoiICML 2026
Builds on16
- TabDDPM: Modelling Tabular Data with Diffusion ModelsAkim Kotelnikov, Dmitry Baranchuk, Ivan Rubachev, Artem BabenkoICML 2023 · 518 citations
- Transformers Can Do Bayesian InferenceSamuel Müller, Noah Hollmann, Sebastian Pineda-Arango, Josif Grabocka et al.ICLR 2022 · 287 citations
- Mixed-Type Tabular Data Synthesis with Score-based Diffusion in Latent SpaceHengrui Zhang, Jiani Zhang, Zhengyuan Shen, Balasubramaniam Srinivasan et al.ICLR 2024 · 233 citations
- Missing Data Imputation using Optimal TransportBoris Muzellec, Julie Josse, Claire Boyer, Marco CuturiICML 2020 · 179 citations
- Masked Language Model ScoringJulian Salazar, Davis Liang, Toan Q. Nguyen, Katrin KirchhoffACL 2020 · 167 citations
Related papers
- TabNAT: A Continuous-Discrete Joint Generative Framework for Tabular DataHengrui Zhang, Liancheng Fang, Qitian Wu, Philip S. YuICML 2025
- TabMT: Generating tabular data with masked transformersManbir S. Gulati, Paul F. RoysdonNeurIPS 2023 · 70 citations
- TabDiff: a Mixed-type Diffusion Model for Tabular Data GenerationJuntong Shi, Minkai Xu, Harper Hua, Hengrui Zhang et al.ICLR 2025
- Language-Interfaced Tabular Oversampling via Progressive Imputation and Self-AuthenticationJune Yong Yang, Geondo Park, Joowon Kim, Hyeongwon Jang et al.ICLR 2024 · 8 citations
- Characterizing the Predictive Impact of Modalities with Supervised Latent-Variable ModelingDivyam Madaan, Sumit Chopra, Kyunghyun ChoICML 2026
