Generalizable Data Cleaning of Tabular Data in Latent Space
Eduardo Souza dos Reis, Mohamed Abdelaal, Carsten Binnig
Abstract
In this paper, we present a new method for learned data cleaning. In contrast to existing methods, our method learns to clean data in the latent space. The main idea is that we (1) shape the latent space such that we know the area where clean data resides and (2) learn latent operators trained on error repair (Lopster) which shift erroneous data (e.g., table rows with noise, outliers, or missing values) in their latent representation back to a "clean" region, thus abstracting the complexities of the input domain. When formulating data cleaning as a simple shift operation in latent space, we can repair all types of errors using the same method which makes it more robust than other methods. Importantly, with our method, we can handle errors that are unseen during the training of our error repair model. We do not rely on an external error detection method as seen in the state-of-the-art, instead, we handle both detection and repair within the Lopster framework. In our evaluation, we show that our approach outperforms existing cleaning methods even when trained on only a subset of the errors that occur in the dirty data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 14c10800-222b-4d38-a091-eec678e1d947Cited by top-tier papers1
Ask how each one uses itBuilds on11
- TURL: Table Understanding through Representation LearningXiang Deng, Huan Sun, Alyssa Lees, You Wu et al.VLDB 2021 · 2,406 citations
- TaBERT: Pretraining for Joint Understanding of Textual and Tabular DataPengcheng Yin, Graham Neubig, Wen-tau Yih, Sebastian RiedelACL 2020 · 417 citations
- Can Foundation Models Wrangle Your Data?Avanika Narayan, Ines Chami, Laurel J. Orr, Christopher RéVLDB 2023 · 325 citations
- Horizon: Scalable Dependency-driven Data CleaningEl Kindi Rezig, Mourad Ouzzani, Walid G. Aref, Ahmed K. Elmagarmid et al.VLDB 2021 · 95 citations
- Missing Value Imputation on Multidimensional Time SeriesParikshit Bansal, Prathamesh Deshpande, Sunita SarawagiVLDB 2021 · 90 citations
Related papers
- DeMix: Debugging Training Data with Mixed Data Error Types by Investigating Influence VectorsJiale Deng, Yanyan Shen, Xiaogang Shi, Junjun ChaiKDD 2026
- TORepair: Diffusion-Based Task-Oriented Error Repair Via Differentiable Bi-Level OptimizationWei Ni, Xiaoye Miao, Xiangyu Zhao, Yangyang Wu et al.ICDE 2026
- A-DARTS: Stable Model Selection for Data Repair in Time SeriesMourad Khayati, Guillaume Chacun, Zakhar Tymchenko, Philippe Cudré-MaurouxICDE 2025 · 2 citations
- A Topological Filter for Learning with Label NoisePengxiang Wu, Songzhu Zheng, Mayank Goswami, Dimitris N. Metaxas et al.NeurIPS 2020 · 143 citations
- CleanML: A Study for Evaluating the Impact of Data Cleaning on ML Classification TasksPeng Li, Xi Rao, Jennifer Blase, Yue Zhang et al.ICDE 2021 · 127 citations
