Generalizable Data Cleaning of Tabular Data in Latent Space
Eduardo Souza dos Reis, Mohamed Abdelaal, Carsten Binnig
摘要
In this paper, we present a new method for learned data cleaning. In contrast to existing methods, our method learns to clean data in the latent space. The main idea is that we (1) shape the latent space such that we know the area where clean data resides and (2) learn latent operators trained on error repair (Lopster) which shift erroneous data (e.g., table rows with noise, outliers, or missing values) in their latent representation back to a "clean" region, thus abstracting the complexities of the input domain. When formulating data cleaning as a simple shift operation in latent space, we can repair all types of errors using the same method which makes it more robust than other methods. Importantly, with our method, we can handle errors that are unseen during the training of our error repair model. We do not rely on an external error detection method as seen in the state-of-the-art, instead, we handle both detection and repair within the Lopster framework. In our evaluation, we show that our approach outperforms existing cleaning methods even when trained on only a subset of the errors that occur in the dirty data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper11
- TURL: Table Understanding through Representation LearningXiang Deng, Huan Sun, Alyssa Lees, You Wu 等VLDB 2021 · 被引用 2,406 次
- TaBERT: Pretraining for Joint Understanding of Textual and Tabular DataPengcheng Yin, Graham Neubig, Wen-tau Yih, Sebastian RiedelACL 2020 · 被引用 417 次
- Can Foundation Models Wrangle Your Data?Avanika Narayan, Ines Chami, Laurel J. Orr, Christopher RéVLDB 2023 · 被引用 325 次
- Horizon: Scalable Dependency-driven Data CleaningEl Kindi Rezig, Mourad Ouzzani, Walid G. Aref, Ahmed K. Elmagarmid 等VLDB 2021 · 被引用 95 次
- Missing Value Imputation on Multidimensional Time SeriesParikshit Bansal, Prathamesh Deshpande, Sunita SarawagiVLDB 2021 · 被引用 90 次
相关 Paper
- DeMix: Debugging Training Data with Mixed Data Error Types by Investigating Influence VectorsJiale Deng, Yanyan Shen, Xiaogang Shi, Junjun ChaiKDD 2026
- TORepair: Diffusion-Based Task-Oriented Error Repair Via Differentiable Bi-Level OptimizationWei Ni, Xiaoye Miao, Xiangyu Zhao, Yangyang Wu 等ICDE 2026
- A-DARTS: Stable Model Selection for Data Repair in Time SeriesMourad Khayati, Guillaume Chacun, Zakhar Tymchenko, Philippe Cudré-MaurouxICDE 2025 · 被引用 2 次
- A Topological Filter for Learning with Label NoisePengxiang Wu, Songzhu Zheng, Mayank Goswami, Dimitris N. Metaxas 等NeurIPS 2020 · 被引用 143 次
- CleanML: A Study for Evaluating the Impact of Data Cleaning on ML Classification TasksPeng Li, Xi Rao, Jennifer Blase, Yue Zhang 等ICDE 2021 · 被引用 127 次
