To Predict or Not to Predict? Proportionally Masked Autoencoders for Tabular Data Imputation
Jungkyu Kim, Kibok Lee, Taeyoung Park
Abstract
Masked autoencoders (MAEs) have recently demonstrated effectiveness in tabular data imputation. However, due to the inherent heterogeneity of tabular data, the uniform random masking strategy commonly used in MAEs can disrupt the distribution of missingness, leading to suboptimal performance. To address this, we propose a proportional masking strategy for MAEs. Specifically, we first compute the statistics of missingness based on the observed proportions in the dataset, and then generate masks that align with these statistics, ensuring that the distribution of missingness is preserved after masking. Furthermore, we argue that simple MLP-based token mixing offers competitive or often superior performance compared to attention mechanisms while being more computationally efficient, especially in the tabular domain with the inherent heterogeneity. Experimental results validate the effectiveness of the proposed proportional masking strategy across various missing data patterns in tabular datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0bfac0d1-8570-404e-80e0-2d313de22212Cited by top-tier papers4
- Hybrid Autoencoders for Tabular Data: Leveraging Model-Based Augmentation in Low-Label SettingsErel Naor, Ofir LindenbaumNeurIPS 2025 · 6 citations
- Impute Missing Entries with UncertaintyJaesung Lim, Seunghwan An, Jong-June JeonAAAI 2026
- Towards Synthesizing High-Dimensional Tabular Data with Limited SamplesZuqing Li, Junhao Gan, Jianzhong QiAAAI 2026
- Deep Incomplete Multi-View Clustering via Hierarchical Imputation and AlignmentYiming Du, Ziyu Wang, Jian Li, Rui Ning et al.AAAI 2026
Builds on13
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
- Mixed-Type Tabular Data Synthesis with Score-based Diffusion in Latent SpaceHengrui Zhang, Jiani Zhang, Zhengyuan Shen, Balasubramaniam Srinivasan et al.ICLR 2024 · 233 citations
- HyperImpute: Generalized Iterative Imputation with Automatic Model SelectionDaniel Jarrett, Bogdan Cebere, Tennison Liu, Alicia Curth et al.ICML 2022 · 129 citations
- MIRACLE: Causally-Aware Imputation via Learning Missing Data MechanismsTrent Kyono, Yao Zhang, Alexis Bellot, Mihaela van der SchaarNeurIPS 2021 · 105 citations
- Enhanced Doubly Robust Learning for Debiasing Post-Click Conversion Rate EstimationSiyuan Guo, Lixin Zou, Yiding Liu, Wenwen Ye et al.SIGIR 2021 · 63 citations
Related papers
- ReMasker: Imputing Tabular Data with Masked AutoencodingTianyu Du, Luca Melis, Ting WangICLR 2024 · 41 citations
- CACTI: Leveraging Copy Masking and Contextual Information to Improve Tabular Data ImputationAditya Gorla, Ryan Wang, Zhengtong Liu, Ulzee An et al.ICML 2025
- TabNAT: A Continuous-Discrete Joint Generative Framework for Tabular DataHengrui Zhang, Liancheng Fang, Qitian Wu, Philip S. YuICML 2025
- MISS: An Incomplete Tabular Data Representation System with Missing Mechanism LearningYangyang Wu, Shuwei Liang, Lei Qiang, Xiaoye Miao et al.ICDE 2025
- Diffusion Transformers for Tabular Data Time Series GenerationFabrizio Garuti, Enver Sangineto, Simone Luetto, Lorenzo Forni et al.ICLR 2025 · 12 citations
