To Predict or Not to Predict? Proportionally Masked Autoencoders for Tabular Data Imputation
Jungkyu Kim, Kibok Lee, Taeyoung Park
摘要
Masked autoencoders (MAEs) have recently demonstrated effectiveness in tabular data imputation. However, due to the inherent heterogeneity of tabular data, the uniform random masking strategy commonly used in MAEs can disrupt the distribution of missingness, leading to suboptimal performance. To address this, we propose a proportional masking strategy for MAEs. Specifically, we first compute the statistics of missingness based on the observed proportions in the dataset, and then generate masks that align with these statistics, ensuring that the distribution of missingness is preserved after masking. Furthermore, we argue that simple MLP-based token mixing offers competitive or often superior performance compared to attention mechanisms while being more computationally efficient, especially in the tabular domain with the inherent heterogeneity. Experimental results validate the effectiveness of the proposed proportional masking strategy across various missing data patterns in tabular datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Hybrid Autoencoders for Tabular Data: Leveraging Model-Based Augmentation in Low-Label SettingsErel Naor, Ofir LindenbaumNeurIPS 2025 · 被引用 6 次
- Impute Missing Entries with UncertaintyJaesung Lim, Seunghwan An, Jong-June JeonAAAI 2026
- Towards Synthesizing High-Dimensional Tabular Data with Limited SamplesZuqing Li, Junhao Gan, Jianzhong QiAAAI 2026
- Deep Incomplete Multi-View Clustering via Hierarchical Imputation and AlignmentYiming Du, Ziyu Wang, Jian Li, Rui Ning 等AAAI 2026
它引用的顶会 Paper13
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer 等NeurIPS 2021 · 被引用 3,862 次
- Mixed-Type Tabular Data Synthesis with Score-based Diffusion in Latent SpaceHengrui Zhang, Jiani Zhang, Zhengyuan Shen, Balasubramaniam Srinivasan 等ICLR 2024 · 被引用 233 次
- HyperImpute: Generalized Iterative Imputation with Automatic Model SelectionDaniel Jarrett, Bogdan Cebere, Tennison Liu, Alicia Curth 等ICML 2022 · 被引用 129 次
- MIRACLE: Causally-Aware Imputation via Learning Missing Data MechanismsTrent Kyono, Yao Zhang, Alexis Bellot, Mihaela van der SchaarNeurIPS 2021 · 被引用 105 次
- Enhanced Doubly Robust Learning for Debiasing Post-Click Conversion Rate EstimationSiyuan Guo, Lixin Zou, Yiding Liu, Wenwen Ye 等SIGIR 2021 · 被引用 63 次
相关 Paper
- ReMasker: Imputing Tabular Data with Masked AutoencodingTianyu Du, Luca Melis, Ting WangICLR 2024 · 被引用 41 次
- CACTI: Leveraging Copy Masking and Contextual Information to Improve Tabular Data ImputationAditya Gorla, Ryan Wang, Zhengtong Liu, Ulzee An 等ICML 2025
- TabNAT: A Continuous-Discrete Joint Generative Framework for Tabular DataHengrui Zhang, Liancheng Fang, Qitian Wu, Philip S. YuICML 2025
- MISS: An Incomplete Tabular Data Representation System with Missing Mechanism LearningYangyang Wu, Shuwei Liang, Lei Qiang, Xiaoye Miao 等ICDE 2025
- Diffusion Transformers for Tabular Data Time Series GenerationFabrizio Garuti, Enver Sangineto, Simone Luetto, Lorenzo Forni 等ICLR 2025 · 被引用 12 次
