Self-supervised and Interpretable Data Cleaning with Sequence Generative Adversarial Networks
Jinfeng Peng, Derong Shen, Nan Tang, Tieying Liu, Yue Kou, Tiezheng Nie, Hang Cui, Ge Yu
摘要
We study the problem of self-supervised and interpretable data cleaning, which automatically extracts interpretable data repair rules from dirty data. In this paper, we propose a novel framework, namely Garf, based on sequence generative adversarial networks (SeqGAN). One key information Garf tries to capture is data repair rules (for example, if the city is "Dothan", then the county should be "Houston"). Garf employs a SeqGAN consisting of a generator G and a discriminator D that trains G to learn the dependency relationships ( e.g. , given a city value "Dothan" as input, the county can be determined as "Houston"). After training, the generator G can be used to generate data repair rules, but may contain both trusted and untrusted rules, especially when learning from dirty data. To mitigate this problem, Garf further updates the learned relationships with another discriminator D' to iteratively improve the quality of both rules and data. Garf takes advantages of both logical and learning-based methods, which allow cleaning dirty data with high interpretability and have no requirements for prior knowledge and training data. Extensive experiments on real-world and synthetic datasets demonstrate the effectiveness of Garf. Garf achieves new state-of-the-art data cleaning result with high accuracy, through learning from dirty datasets without human supervision.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- TSM-Bench: Benchmarking Time Series Database Systems for Monitoring ApplicationsAbdelouahab Khelifati, Mourad Khayati, Anton Dignös, Djellel Eddine Difallah 等VLDB 2023 · 被引用 24 次
- In-Database Data ImputationMassimo Perini, Milos NikolicSIGMOD 2024 · 被引用 5 次
- BClean: A Bayesian Data Cleaning SystemJianbin Qin, Sifan Huang, Yaoshu Wang, Jing Zhu 等ICDE 2024 · 被引用 4 次
- Generalizable Data Cleaning of Tabular Data in Latent SpaceEduardo Souza dos Reis, Mohamed Abdelaal, Carsten BinnigVLDB 2024 · 被引用 3 次
- Mitigating Data Sparsity in Integrated Data through Text ConceptualizationMd. Ataur Rahman, Sergi Nadal, Oscar Romero, Dimitris SacharidisICDE 2024 · 被引用 3 次
它引用的顶会 Paper7
- TURL: Table Understanding through Representation LearningXiang Deng, Huan Sun, Alyssa Lees, You Wu 等VLDB 2021 · 被引用 2,406 次
- Free-Form Image Inpainting With Gated ConvolutionJiahui Yu, Zhe Lin, Jimei Yang, Xiaohui Shen 等ICCV 2019 · 被引用 1,990 次
- Horizon: Scalable Dependency-driven Data CleaningEl Kindi Rezig, Mourad Ouzzani, Walid G. Aref, Ahmed K. Elmagarmid 等VLDB 2021 · 被引用 95 次
- RPT: Relational Pre-trained Transformer Is Almost All You Need towards Democratizing Data PreparationNan Tang, Ju Fan, Fangyi Li, Jianhong Tu 等VLDB 2021 · 被引用 92 次
- GAN Ensemble for Anomaly DetectionXu Han, Xiaohui Chen, Li-Ping LiuAAAI 2021 · 被引用 79 次
相关 Paper
- GIDCL: A Graph-Enhanced Interpretable Data Cleaning Framework with Large Language ModelsMengyi Yan, Yaoshu Wang, Yue Wang, Xiaoye Miao 等SIGMOD 2025 · 被引用 12 次
- Self-supervised Adversarial Purification for Graph Neural NetworksWoohyun Lee, Hogun ParkICML 2025
- Adaptive Data Augmentation for Supervised Learning over Missing DataTongyu Liu, Ju Fan, Yinqing Luo, Nan Tang 等VLDB 2021 · 被引用 31 次
- Interpretable Generative Adversarial NetworksChao Li, Kelu Yao, Jin Wang, Boyu Diao 等AAAI 2022 · 被引用 19 次
- ARGAN: Attentive Recurrent Generative Adversarial Network for Shadow Detection and RemovalBin Ding, Chengjiang Long, Ling Zhang, Chunxia XiaoICCV 2019 · 被引用 171 次
