ZeroED: Hybrid Zero-Shot Error Detection Through Large Language Model Reasoning
Wei Ni, Kaihang Zhang, Xiaoye Miao, Xiangyu Zhao, Yangyang Wu, Yaoshu Wang, Jianwei Yin
Abstract
Error detection (ED) in tabular data is crucial yet challenging due to diverse error types and the need for contextual understanding. Traditional ED methods often rely heavily on manual criteria and labels, making them labor-intensive. Large language models (LLM) can minimize human effort but struggle with errors requiring a comprehensive understanding of data context. In this paper, we propose ZeroED, a novel hybrid error detection framework, which combines LLM reasoning ability with the machine learning pipeline via zero-shot prompting. ZeroED operates in four steps, i.e., feature representation, error labeling, training data construction, and detector training. Initially, to enhance error distinction, ZeroED generates rich data representations using LLM-driven error reason-aware binary features, pre-trained embeddings, and statistical features. Then, ZeroED employs LLM to holistically label errors through incontext learning, guided by a two-step LLM reasoning process for detailed ED guidelines. To reduce token costs, LLMs are applied only to representative data selected via clustering-based sampling. High-quality training data is constructed through in-cluster label propagation and LLM augmentation with verification. Finally, a classifier is trained to detect all errors. Extensive experiments on seven datasets demonstrate that, ZeroED outperforms state-of-the-art methods by a maximum 30 % improvement in F1 score and up to 90% token cost reduction.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7d444ec7-a2ff-4f88-baa5-4e3ef444df66Cited by top-tier papers3
- Stepwise Reasoning Disruption Attack of LLMsJingyu Peng, Maolin Wang, Xiangyu Zhao, Kai Zhang et al.ACL 2025 · 11 citations
- Ensembling LLM-Induced Decision Trees for Explainable and Robust Error DetectionMengqi Wang, Jianwei Wang, Qing Liu, Xiwei Xu et al.KDD 2026 · 4 citations
- Empowering Tabular Data Preparation with Language Models: Why and How?Mengshi Chen, Yuxiang Sun, Tengchao Li, Jianwei Wang et al.ACL 2026 · 4 citations
Builds on12
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
Related papers
- Label Annotation for Tabular Anomaly Detection with Large Language ModelsHaihong Zhao, Aochuan Chen, Miao Peng, Xiaolong Fan et al.KDD 2026
- Large Language Models Can Automatically Engineer Features for Few-Shot Tabular LearningSungwon Han, Jinsung Yoon, Sercan Ö. Arik, Tomas PfisterICML 2024 · 81 citations
- Are LLMs Good Zero-Shot Fallacy Classifiers?Fengjun Pan, Xiaobao Wu, Zongrui Li, Anh Tuan LuuEMNLP 2024 · 7 citations
- ZeroRel: Relational Reasoning via Graph-guided Large Language ModelsYujie Tian, Kun Zhang, Qiuyu Li, Le Wu et al.KDD 2026
- ReTabAD: A Benchmark for Restoring Semantic Context in Tabular Anomaly DetectionSanghyu Yoon, Dongmin Kim, Suhee Yoon, Ye Seul Sim et al.ICLR 2026 · 3 citations
