A Zero-Training Error Correction System with Large Language Models
Yangyang Wu, Chen Yang, Mengying Zhu, Xiaoye Miao, Wei Ni, Meng Xi, Xinkui Zhao, Jianwei Yin
摘要
Correcting missing or erroneous data values is an essential task in data cleaning. Traditional pre-configuration error correction (EC) methods rely heavily on predefined rules or constraints, demanding significant domain knowledge and manual effort. While configuration-free EC approaches have been explored, they still demand extensive feature engineering or labeled data for intensive model training. In this paper, we propose a zero-training and interpretable EC system, named ZeroEC, that leverages large language models (LLMs) to generate chain-of-thoughts (CoTs) and correction rules for EC, without the need for model training. ZeroEC consists of two modules, contextual-relevant tuple search (CTS) and training-free explainable correction (TEC). CTS constructs a contextual-relevant tuple retriever using a weighted cosine similarity function to efficiently identify the most relevant tuples for each dirty tuple, reducing redundancy in the LLM prompts and lowering computational costs. TEC employs a clustering-based representative tuple sampling strategy to alleviate “hallucination” risk by exposing LLMs to diverse types of data errors. It further prompts for generating correction CoTs for user-corrected representative tuples, as well as prompts for creating correction rules and explainable ECs, which automatically provide explanations for EC, all without the need for model training. Extensive experiments conducted on various real-world datasets demonstrate that ZeroEC achieves a 66.82% increase in accuracy and a 6.87x speedup compared to state-of-the-art methods. The codes and datasets of this paper are available at https://github.com/YangChen32768/ZeroEC.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- GIDCL: A Graph-Enhanced Interpretable Data Cleaning Framework with Large Language ModelsMengyi Yan, Yaoshu Wang, Yue Wang, Xiaoye Miao 等SIGMOD 2025 · 被引用 12 次
- A Training-free LLM-based Approach to General Chinese Character Error CorrectionHouquan Zhou, Bo Zhang, Zhenghua Li, Ming Yan 等ACL 2025
- ZeroED: Hybrid Zero-Shot Error Detection Through Large Language Model ReasoningWei Ni, Kaihang Zhang, Xiaoye Miao, Xiangyu Zhao 等ICDE 2025 · 被引用 5 次
- Learning by Correction: Efficient Tuning Task for Zero-Shot Generative Vision-Language ReasoningRongjie Li, Yu Wu, Xuming HeCVPR 2024
- CEC-Zero: Zero-Supervision Character Error Correction with Self-Generated RewardsZhiming Lin, Kai Zhao, Sophie Zhang, Peilai Yu 等AAAI 2026 · 被引用 11 次
