CACTI: Leveraging Copy Masking and Contextual Information to Improve Tabular Data Imputation
Aditya Gorla, Ryan Wang, Zhengtong Liu, Ulzee An, Sriram Sankararaman
Abstract
We present CACTI, a masked autoencoding approach for imputing tabular data that leverages the structure in missingness patterns and contextual information. Our approach employs a novel median truncated copy masking training strategy that encourages the model to learn from empirical patterns of missingness while incorporating semantic relationships between features -captured by column names and text descriptions -to better represent feature dependence. These dual sources of inductive bias enable CACTI to outperform state-of-the-art methods -an average R 2 gain of 7.8% over the next best method (13.4%, 6.1%, and 5.3% under missing not at random, at random and completely at random, respectively)across a diverse range of datasets and missingness conditions. Our results highlight the value of leveraging dataset-specific contextual information and missingness patterns to enhance imputation performance. Code is publicly available at github.com/sriramlab/CACTI A primary reason underlying this challenge is that missing-
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on9
- TabNet: Attentive Interpretable Tabular LearningSercan Ö. Arik, Tomas PfisterAAAI 2021 · 2,148 citations
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 394 citations
- Missing Data Imputation using Optimal TransportBoris Muzellec, Julie Josse, Claire Boyer, Marco CuturiICML 2020 · 179 citations
- Large Scale Transfer Learning for Tabular Data via Language ModelingJosh Gardner, Juan C. Perdomo, Ludwig SchmidtNeurIPS 2024 · 103 citations
- ReMasker: Imputing Tabular Data with Masked AutoencodingTianyu Du, Luca Melis, Ting WangICLR 2024 · 41 citations
Related papers
- To Predict or Not to Predict? Proportionally Masked Autoencoders for Tabular Data ImputationJungkyu Kim, Kibok Lee, Taeyoung ParkAAAI 2025 · 4 citations
- MIRACLE: Causally-Aware Imputation via Learning Missing Data MechanismsTrent Kyono, Yao Zhang, Alexis Bellot, Mihaela van der SchaarNeurIPS 2021 · 105 citations
- MISS: An Incomplete Tabular Data Representation System with Missing Mechanism LearningYangyang Wu, Shuwei Liang, Lei Qiang, Xiaoye Miao et al.ICDE 2025
- CaT-Diff: Cascaded Text-enhanced Diffusion Model for Time-Series ImputationChangjian Xu, Yong Wang, Ruizheng Huang, Zhicheng Zhang et al.AAAI 2026
- Missing Data Imputation and Acquisition with Deep Hierarchical Models and Hamiltonian Monte CarloIgnacio Peis, Chao Ma, José Miguel Hernández-LobatoNeurIPS 2022 · 25 citations
