GIDCL: A Graph-Enhanced Interpretable Data Cleaning Framework with Large Language Models
Mengyi Yan, Yaoshu Wang, Yue Wang, Xiaoye Miao, Jianxin Li
Abstract
Data quality is critical across many applications. The utility of data is undermined by various errors, making rigorous data cleaning a necessity. Traditional data cleaning systems depend heavily on predefined rules and constraints, which necessitate significant domain knowledge and manual effort. Moreover, while configuration-free approaches and deep learning methods have been explored, they struggle with complex error patterns, lacking interpretability, requiring extensive feature engineering or labeled data. This paper introduces GIDCL ( G raph-enhanced I nterpretable D ata C leaning with L arge language models), a pioneering framework that harnesses the capabilities of Large Language Models (LLMs) alongside Graph Neural Network (GNN) to address the challenges of traditional and machine learning-based data cleaning methods. By converting relational tables into graph structures, GIDCL utilizes GNN to effectively capture and leverage structural correlations among data, enhancing the model's ability to understand and rectify complex dependencies and errors. The framework's creator-critic workflow innovatively employs LLMs to automatically generate interpretable data cleaning rules and tailor feature engineering with minimal labeled data. This process includes the iterative refinement of error detection and correction models through few-shot learning, significantly reducing the need for extensive manual configuration. GIDCL not only improves the precision and efficiency of data cleaning but also enhances its interpretability, making it accessible and practical for non-expert users. Our extensive experiments demonstrate that GIDCL significantly outperforms existing methods, improving F1-scores by 10% on average while requiring only 20 labeled tuples.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 63b848f9-262a-47b8-9403-c5a8620174ffCited by top-tier papers4
- This is Going to Sound Crazy, But What If We Used Large Language Models to Boost Automatic Database Tuning Algorithms By Leveraging Prior History? We Will Find Better Configurations More Quickly Than Retraining From Scratch!William Zhang, Wan Shen Lim, Andrew PavloSIGMOD 2026 · 7 citations
- CatDB: Data-catalog-guided, LLM-based Generation of Data-centric ML PipelinesSaeed Fathollahzadeh, Essam Mansour, Matthias BoehmVLDB 2025 · 5 citations
- Empowering Tabular Data Preparation with Language Models: Why and How?Mengshi Chen, Yuxiang Sun, Tengchao Li, Jianwei Wang et al.ACL 2026 · 4 citations
- Accurate Table Question Answering with Accessible LLMsYangfan Jiang, Fei Wei, Ergute Bao, Yaliang Li et al.ICDE 2026 · 1 citation
Related papers
- A Zero-Training Error Correction System with Large Language ModelsYangyang Wu, Chen Yang, Mengying Zhu, Xiaoye Miao et al.ICDE 2025 · 7 citations
- Can LLMs Find Fraudsters? Multi-level LLM Enhanced Graph Fraud DetectionTairan Huang, Yili Wang, Qiutong Li, Changlong He et al.ACM MM 2025 · 10 citations
- Self-supervised and Interpretable Data Cleaning with Sequence Generative Adversarial NetworksJinfeng Peng, Derong Shen, Nan Tang, Tieying Liu et al.VLDB 2023 · 24 citations
- ZeroED: Hybrid Zero-Shot Error Detection Through Large Language Model ReasoningWei Ni, Kaihang Zhang, Xiaoye Miao, Xiangyu Zhao et al.ICDE 2025 · 5 citations
- Bootstrapping Heterogeneous Graph Representation Learning via Large Language Models: A Generalized ApproachHang Gao, Chenhao Zhang, Fengge Wu, Changwen Zheng et al.AAAI 2025 · 6 citations
