MISS: An Incomplete Tabular Data Representation System with Missing Mechanism Learning
Yangyang Wu, Shuwei Liang, Lei Qiang, Xiaoye Miao, Xinkui Zhao, Junlan Cai, Yunjun Gao, Jianwei Yin
Abstract
The missing data problem widely exists in real-life scenarios. The incomplete data analysis through imputation can amplify the errors or bias, hindering the effective analysis. Ex-isting tabular data representation methods overlook the missing state of data values, and thus cannot effectively deal with the incomplete data. In this paper, we propose a novel incomplete tabular data representation system, named MISS. It is capable of enabling all Transformer-based tabular representation methods to effectively handle incomplete data. MISS consists of two modules, i.e., missing mechanism learning (MML) and incomplete data representation (IDR). MML leverages a new missingness propensity score calculation strategy to learn the observed data distribution and missing mechanisms within incomplete data. IDR introduces a novel probability-driven Transformer block, in conjunction with an unbiased representation loss function, for effective representation. We prove that, MISS can eliminate the bias resulting from missingness. Extensive experiments on four public real-world datasets demonstrate that, MISS yields a more than 57 % accuracy gain with competitive efficiency, compared with the state-of-the-art approaches.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 27b6a3e2-d31e-4cb8-9c7e-b4f7fd8945e8Related papers
- Fairness-Aware Classification over Incomplete DataXiaoye Miao, Lei Qiang, Guilin Huang, Yangyang Wu et al.SIGIR 2025 · 1 citation
- ReMasker: Imputing Tabular Data with Masked AutoencodingTianyu Du, Luca Melis, Ting WangICLR 2024 · 41 citations
- Federated Incomplete Tabular Data Prediction with Missing ComplementarityYan Zhang, Shuwei Liang, Xiaoye Miao, Yangyang Wu et al.VLDB 2025
- To Predict or Not to Predict? Proportionally Masked Autoencoders for Tabular Data ImputationJungkyu Kim, Kibok Lee, Taeyoung ParkAAAI 2025 · 4 citations
- A Transformer-based Framework for Multivariate Time Series Representation LearningGeorge Zerveas, Srideepika Jayaraman, Dhaval Patel, Anuradha Bhamidipaty et al.KDD 2021 · 66 citations
