ReMasker: Imputing Tabular Data with Masked Autoencoding
Tianyu Du, Luca Melis, Ting Wang
Abstract
We present ReMasker, a new method of imputing missing values in tabular data by extending the masked autoencoding framework. Compared with prior work, ReMasker is both simple -- besides the missing values (i.e., naturally masked), we randomly ``re-mask'' another set of values, optimize the autoencoder by reconstructing this re-masked set, and apply the trained model to predict the missing values; and effective -- with extensive evaluation on benchmark datasets, we show that ReMasker performs on par with or outperforms state-of-the-art methods in terms of both imputation fidelity and utility under various missingness settings, while its performance advantage often increases with the ratio of missing data. We further explore theoretical justification for its effectiveness, showing that ReMasker tends to learn missingness-invariant representations of tabular data. Our findings indicate that masked modeling represents a promising direction for further research on tabular data imputation. The code is publicly available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5ffdcbe8-757b-427b-8456-267d99f0139eCited by top-tier papers13
- Rethinking the Diffusion Models for Missing Data Imputation: A Gradient Flow PerspectiveZhichao Chen, Haoxuan Li, Fangyikang Wang, Odin Zhang et al.NeurIPS 2024 · 38 citations
- On LLM-Enhanced Mixed-Type Data Imputation with High-Order Message PassingJianmin Wang, Kai Wang, Ying Zhang, Wenjie Zhang et al.VLDB 2025 · 15 citations
- Iterative Missing Data Imputation with Model Form Adaptation and Non-Missing Feature SupervisionHao Wang, Zhengnan Li, Zhichao Chen, Xu Chen et al.NeurIPS 2025 · 12 citations
- To Predict or Not to Predict? Proportionally Masked Autoencoders for Tabular Data ImputationJungkyu Kim, Kibok Lee, Taeyoung ParkAAAI 2025 · 4 citations
- Masked Language Modeling Becomes Conditional Density Estimation for Tabular Data SynthesisSeunghwan An, Gyeongdong Woo, Jaesung Lim, Chang-Hyun Kim et al.AAAI 2025 · 2 citations
Builds on13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
- TransGAN: Two Pure Transformers Can Make One Strong GAN, and That Can Scale UpYifan Jiang, Shiyu Chang, Zhangyang WangNeurIPS 2021 · 515 citations
- On Embeddings for Numerical Features in Tabular Deep LearningYury Gorishniy, Ivan Rubachev, Artem BabenkoNeurIPS 2022 · 338 citations
- StyleSwin: Transformer-based GAN for High-resolution Image GenerationBowen Zhang, Shuyang Gu, Bo Zhang, Jianmin Bao et al.CVPR 2022 · 217 citations
Related papers
- CACTI: Leveraging Copy Masking and Contextual Information to Improve Tabular Data ImputationAditya Gorla, Ryan Wang, Zhengtong Liu, Ulzee An et al.ICML 2025
- AugMask: Training Diffusion Models on Incomplete Tabular Data via Stochastic Augmentation and MaskingJungkyu Kim, Taeyoung Park, Kibok LeeICML 2026
- TabNAT: A Continuous-Discrete Joint Generative Framework for Tabular DataHengrui Zhang, Liancheng Fang, Qitian Wu, Philip S. YuICML 2025
- MISS: An Incomplete Tabular Data Representation System with Missing Mechanism LearningYangyang Wu, Shuwei Liang, Lei Qiang, Xiaoye Miao et al.ICDE 2025
- TabMT: Generating tabular data with masked transformersManbir S. Gulati, Paul F. RoysdonNeurIPS 2023 · 70 citations
