TabMT: Generating tabular data with masked transformers
Manbir S. Gulati, Paul F. Roysdon
Abstract
Autoregressive and Masked Transformers are incredibly effective as generative models and classifiers. While these models are most prevalent in NLP, they also exhibit strong performance in other domains, such as vision. This work contributes to the exploration of transformer-based models in synthetic data generation for diverse application domains. In this paper, we present TabMT, a novel Masked Transformer design for generating synthetic tabular data. TabMT effectively addresses the unique challenges posed by heterogeneous data fields and is natively able to handle missing data. Our design leverages improved masking techniques to allow for generation and demonstrates state-of-the-art performance from extremely small to extremely large tabular datasets. We evaluate TabMT for privacy-focused applications and find that it is able to generate high quality data with superior privacy tradeoffs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1ceeb488-46f3-4fa2-aae6-fb777084afc2Cited by top-tier papers11
- Diffusion Transformers for Tabular Data Time Series GenerationFabrizio Garuti, Enver Sangineto, Simone Luetto, Lorenzo Forni et al.ICLR 2025 · 12 citations
- TabStruct: Measuring Structural Fidelity of Tabular DataXiangjian Jiang, Nikola Simidjievski, Mateja JamnikICLR 2026 · 10 citations
- Large Language Models are Good Relational LearnersFang Wu, Vijay Prakash Dwivedi, Jure LeskovecACL 2025 · 10 citations
- MUSE: Model-Agnostic Tabular Watermarking via Multi-Sample SelectionLiancheng Fang, Aiwei Liu, Henry Peng Zou, Yankai Chen et al.ICLR 2026 · 5 citations
- Table Question Answering in the Era of Large Language Models: A Comprehensive Survey of Tasks, Methods, and EvaluationWei Zhou, Bolei Ma, Annemarie Friedrich, Mohsen MesgarACL 2026 · 3 citations
Builds on9
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu et al.ICML 2020 · 1,773 citations
- Masked Autoencoders that ListenPo-Yao Huang, Hu Xu, Juncheng Li, Alexei Baevski et al.NeurIPS 2022 · 524 citations
- TabDDPM: Modelling Tabular Data with Diffusion ModelsAkim Kotelnikov, Dmitry Baranchuk, Ivan Rubachev, Artem BabenkoICML 2023 · 518 citations
Related papers
- Language Models are Realistic Tabular Data GeneratorsVadim Borisov, Kathrin Seßler, Tobias Leemann, Martin Pawelczyk et al.ICLR 2023 · 45 citations
- TabNAT: A Continuous-Discrete Joint Generative Framework for Tabular DataHengrui Zhang, Liancheng Fang, Qitian Wu, Philip S. YuICML 2025
- MultiTab: A Scalable Foundation for Multitask Learning on Tabular DataDimitrios Sinodinos, Jack Yi Wei, Narges ArmanfardAAAI 2026 · 2 citations
- Generative Table Pre-training Empowers Models for Tabular PredictionTianping Zhang, Shaowen Wang, Shuicheng Yan, Li Jian et al.EMNLP 2023 · 18 citations
- Robust Detection of Synthetic Tabular Data Under Schema VariabilityG. Charbel N. Kindji, Elisa Fromont, Lina Maria Rojas-Barahona, Tanguy UrvoyAAAI 2026
