Self-supervised and Interpretable Data Cleaning with Sequence Generative Adversarial Networks
Jinfeng Peng, Derong Shen, Nan Tang, Tieying Liu, Yue Kou, Tiezheng Nie, Hang Cui, Ge Yu
Abstract
We study the problem of self-supervised and interpretable data cleaning, which automatically extracts interpretable data repair rules from dirty data. In this paper, we propose a novel framework, namely Garf, based on sequence generative adversarial networks (SeqGAN). One key information Garf tries to capture is data repair rules (for example, if the city is "Dothan", then the county should be "Houston"). Garf employs a SeqGAN consisting of a generator G and a discriminator D that trains G to learn the dependency relationships ( e.g. , given a city value "Dothan" as input, the county can be determined as "Houston"). After training, the generator G can be used to generate data repair rules, but may contain both trusted and untrusted rules, especially when learning from dirty data. To mitigate this problem, Garf further updates the learned relationships with another discriminator D' to iteratively improve the quality of both rules and data. Garf takes advantages of both logical and learning-based methods, which allow cleaning dirty data with high interpretability and have no requirements for prior knowledge and training data. Extensive experiments on real-world and synthetic datasets demonstrate the effectiveness of Garf. Garf achieves new state-of-the-art data cleaning result with high accuracy, through learning from dirty datasets without human supervision.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- TSM-Bench: Benchmarking Time Series Database Systems for Monitoring ApplicationsAbdelouahab Khelifati, Mourad Khayati, Anton Dignös, Djellel Eddine Difallah et al.VLDB 2023 · 24 citations
- In-Database Data ImputationMassimo Perini, Milos NikolicSIGMOD 2024 · 5 citations
- BClean: A Bayesian Data Cleaning SystemJianbin Qin, Sifan Huang, Yaoshu Wang, Jing Zhu et al.ICDE 2024 · 4 citations
- Generalizable Data Cleaning of Tabular Data in Latent SpaceEduardo Souza dos Reis, Mohamed Abdelaal, Carsten BinnigVLDB 2024 · 3 citations
- Mitigating Data Sparsity in Integrated Data through Text ConceptualizationMd. Ataur Rahman, Sergi Nadal, Oscar Romero, Dimitris SacharidisICDE 2024 · 3 citations
Builds on7
- TURL: Table Understanding through Representation LearningXiang Deng, Huan Sun, Alyssa Lees, You Wu et al.VLDB 2021 · 2,406 citations
- Free-Form Image Inpainting With Gated ConvolutionJiahui Yu, Zhe Lin, Jimei Yang, Xiaohui Shen et al.ICCV 2019 · 1,990 citations
- Horizon: Scalable Dependency-driven Data CleaningEl Kindi Rezig, Mourad Ouzzani, Walid G. Aref, Ahmed K. Elmagarmid et al.VLDB 2021 · 95 citations
- RPT: Relational Pre-trained Transformer Is Almost All You Need towards Democratizing Data PreparationNan Tang, Ju Fan, Fangyi Li, Jianhong Tu et al.VLDB 2021 · 92 citations
- GAN Ensemble for Anomaly DetectionXu Han, Xiaohui Chen, Li-Ping LiuAAAI 2021 · 79 citations
Related papers
- GIDCL: A Graph-Enhanced Interpretable Data Cleaning Framework with Large Language ModelsMengyi Yan, Yaoshu Wang, Yue Wang, Xiaoye Miao et al.SIGMOD 2025 · 12 citations
- Self-supervised Adversarial Purification for Graph Neural NetworksWoohyun Lee, Hogun ParkICML 2025
- Adaptive Data Augmentation for Supervised Learning over Missing DataTongyu Liu, Ju Fan, Yinqing Luo, Nan Tang et al.VLDB 2021 · 31 citations
- Interpretable Generative Adversarial NetworksChao Li, Kelu Yao, Jin Wang, Boyu Diao et al.AAAI 2022 · 19 citations
- ARGAN: Attentive Recurrent Generative Adversarial Network for Shadow Detection and RemovalBin Ding, Chengjiang Long, Ling Zhang, Chunxia XiaoICCV 2019 · 171 citations
