Masked Language Modeling Becomes Conditional Density Estimation for Tabular Data Synthesis
Seunghwan An, Gyeongdong Woo, Jaesung Lim, Chang-Hyun Kim, Sungchul Hong, Jong-June Jeon
摘要
In this paper, our goal is to generate synthetic data for heterogeneous (mixed-type) tabular datasets with high machine learning utility (MLu). Since the MLu performance depends on accurately approximating the conditional distributions, we focus on devising a synthetic data generation method based on conditional distribution estimation. We introduce MaCoDE by redefining the consecutive multi-class classification task of Masked Language Modeling (MLM) as histogram-based non-parametric conditional density estimation. Our approach enables the estimation of conditional densities across arbitrary combinations of target and conditional variables. We bridge the theoretical gap between distributional learning and MLM by demonstrating that minimizing the orderless multi-class classification loss leads to minimizing the total variation distance between conditional distributions. To validate our proposed model, we evaluate its performance in synthetic data generation across 10 real-world datasets, demonstrating its ability to adjust data privacy levels easily without re-training. Additionally, since masked input tokens in MLM are analogous to missing data, we further assess its effectiveness in handling training datasets with missing values, including multiple imputations of the missing entries.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Impute Missing Entries with UncertaintyJaesung Lim, Seunghwan An, Jong-June JeonAAAI 2026
- Towards Synthesizing High-Dimensional Tabular Data with Limited SamplesZuqing Li, Junhao Gan, Jianzhong QiAAAI 2026
- TabularBERT: Binning-Based Self-Supervised Learning for Tabular RepresentationBeomjin Park, Seunghwan An, Sungchul Hong, Hosik ChoiICML 2026
它引用的顶会 Paper16
- TabDDPM: Modelling Tabular Data with Diffusion ModelsAkim Kotelnikov, Dmitry Baranchuk, Ivan Rubachev, Artem BabenkoICML 2023 · 被引用 518 次
- Transformers Can Do Bayesian InferenceSamuel Müller, Noah Hollmann, Sebastian Pineda-Arango, Josif Grabocka 等ICLR 2022 · 被引用 287 次
- Mixed-Type Tabular Data Synthesis with Score-based Diffusion in Latent SpaceHengrui Zhang, Jiani Zhang, Zhengyuan Shen, Balasubramaniam Srinivasan 等ICLR 2024 · 被引用 233 次
- Missing Data Imputation using Optimal TransportBoris Muzellec, Julie Josse, Claire Boyer, Marco CuturiICML 2020 · 被引用 179 次
- Masked Language Model ScoringJulian Salazar, Davis Liang, Toan Q. Nguyen, Katrin KirchhoffACL 2020 · 被引用 167 次
相关 Paper
- TabNAT: A Continuous-Discrete Joint Generative Framework for Tabular DataHengrui Zhang, Liancheng Fang, Qitian Wu, Philip S. YuICML 2025
- TabMT: Generating tabular data with masked transformersManbir S. Gulati, Paul F. RoysdonNeurIPS 2023 · 被引用 70 次
- TabDiff: a Mixed-type Diffusion Model for Tabular Data GenerationJuntong Shi, Minkai Xu, Harper Hua, Hengrui Zhang 等ICLR 2025
- Language-Interfaced Tabular Oversampling via Progressive Imputation and Self-AuthenticationJune Yong Yang, Geondo Park, Joowon Kim, Hyeongwon Jang 等ICLR 2024 · 被引用 8 次
- Characterizing the Predictive Impact of Modalities with Supervised Latent-Variable ModelingDivyam Madaan, Sumit Chopra, Kyunghyun ChoICML 2026
