Certain and Approximately Certain Models for Statistical Learning
Cheng Zhen, Nischal Aryal, Arash Termehchy, Amandeep Singh Chabada
Abstract
Real-world data is often incomplete and contains missing values. To train accurate models over real-world datasets, users need to spend a substantial amount of time and resources imputing and finding proper values for missing data items. In this paper, we demonstrate that it is possible to learn accurate models directly from data with missing values for certain training data and target models. We propose a unified approach for checking the necessity of data imputation to learn accurate models across various widely-used machine learning paradigms. We build efficient algorithms with theoretical guarantees to check this necessity and return accurate models in cases where imputation is unnecessary. Our extensive experiments indicate that our proposed algorithms significantly reduce the amount of time and effort needed for data imputation without imposing considerable computational overhead.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on6
- What's a good imputation to predict with missing values?Marine Le Morvan, Julie Josse, Erwan Scornet, Gaël VaroquauxNeurIPS 2021 · 95 citations
- Nearest Neighbor Classifiers over Incomplete Information: From Certain Answers to Certain PredictionsBojan Karlas, Peng Li, Renzhi Wu, Nezihe Merve Gürel et al.VLDB 2021 · 69 citations
- GoodCore: Data-effective and Data-efficient Machine Learning through Coreset Selection over Incomplete DataChengliang Chai, Jiabin Liu, Nan Tang, Ju Fan et al.SIGMOD 2023 · 37 citations
- Adaptive Data Augmentation for Supervised Learning over Missing DataTongyu Liu, Ju Fan, Yinqing Luo, Nan Tang et al.VLDB 2021 · 31 citations
- Proving data-poisoning robustness in decision treesSamuel Drews, Aws Albarghouthi, Loris D'AntoniPLDI 2020 · 19 citations
Related papers
- Imputation for prediction: beware of diminishing returnsMarine Le Morvan, Gaël VaroquauxICLR 2025
- Fairness without Imputation: A Decision Tree Approach for Fair Prediction with Missing ValuesHaewon Jeong, Hao Wang, Flávio P. CalmonAAAI 2022 · 48 citations
- Prediction models that learn to avoid missing valuesLena Stempfle, Anton Matsson, Newton Mwai Kinyanjui, Fredrik D. JohanssonICML 2025
- DIM-SUM: Dynamic IMputation for Smart Utility ManagementRyan Hildebrant, Rahul Atul Bhope, Sharad Mehrotra, Christopher Tull et al.VLDB 2025 · 2 citations
- Handling Missing Data with Graph Representation LearningJiaxuan You, Xiaobai Ma, Daisy Yi Ding, Mykel J. Kochenderfer et al.NeurIPS 2020 · 274 citations
