ReMasker: Imputing Tabular Data with Masked Autoencoding
Tianyu Du, Luca Melis, Ting Wang
摘要
We present ReMasker, a new method of imputing missing values in tabular data by extending the masked autoencoding framework. Compared with prior work, ReMasker is both simple -- besides the missing values (i.e., naturally masked), we randomly ``re-mask'' another set of values, optimize the autoencoder by reconstructing this re-masked set, and apply the trained model to predict the missing values; and effective -- with extensive evaluation on benchmark datasets, we show that ReMasker performs on par with or outperforms state-of-the-art methods in terms of both imputation fidelity and utility under various missingness settings, while its performance advantage often increases with the ratio of missing data. We further explore theoretical justification for its effectiveness, showing that ReMasker tends to learn missingness-invariant representations of tabular data. Our findings indicate that masked modeling represents a promising direction for further research on tabular data imputation. The code is publicly available.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Rethinking the Diffusion Models for Missing Data Imputation: A Gradient Flow PerspectiveZhichao Chen, Haoxuan Li, Fangyikang Wang, Odin Zhang 等NeurIPS 2024 · 被引用 38 次
- On LLM-Enhanced Mixed-Type Data Imputation with High-Order Message PassingJianmin Wang, Kai Wang, Ying Zhang, Wenjie Zhang 等VLDB 2025 · 被引用 15 次
- Iterative Missing Data Imputation with Model Form Adaptation and Non-Missing Feature SupervisionHao Wang, Zhengnan Li, Zhichao Chen, Xu Chen 等NeurIPS 2025 · 被引用 12 次
- To Predict or Not to Predict? Proportionally Masked Autoencoders for Tabular Data ImputationJungkyu Kim, Kibok Lee, Taeyoung ParkAAAI 2025 · 被引用 4 次
- Masked Language Modeling Becomes Conditional Density Estimation for Tabular Data SynthesisSeunghwan An, Gyeongdong Woo, Jaesung Lim, Chang-Hyun Kim 等AAAI 2025 · 被引用 2 次
它引用的顶会 Paper13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 被引用 3,632 次
- TransGAN: Two Pure Transformers Can Make One Strong GAN, and That Can Scale UpYifan Jiang, Shiyu Chang, Zhangyang WangNeurIPS 2021 · 被引用 515 次
- On Embeddings for Numerical Features in Tabular Deep LearningYury Gorishniy, Ivan Rubachev, Artem BabenkoNeurIPS 2022 · 被引用 338 次
- StyleSwin: Transformer-based GAN for High-resolution Image GenerationBowen Zhang, Shuyang Gu, Bo Zhang, Jianmin Bao 等CVPR 2022 · 被引用 217 次
相关 Paper
- CACTI: Leveraging Copy Masking and Contextual Information to Improve Tabular Data ImputationAditya Gorla, Ryan Wang, Zhengtong Liu, Ulzee An 等ICML 2025
- AugMask: Training Diffusion Models on Incomplete Tabular Data via Stochastic Augmentation and MaskingJungkyu Kim, Taeyoung Park, Kibok LeeICML 2026
- TabNAT: A Continuous-Discrete Joint Generative Framework for Tabular DataHengrui Zhang, Liancheng Fang, Qitian Wu, Philip S. YuICML 2025
- MISS: An Incomplete Tabular Data Representation System with Missing Mechanism LearningYangyang Wu, Shuwei Liang, Lei Qiang, Xiaoye Miao 等ICDE 2025
- TabMT: Generating tabular data with masked transformersManbir S. Gulati, Paul F. RoysdonNeurIPS 2023 · 被引用 70 次
