Auto-Test: Learning Semantic-Domain Constraints for Unsupervised Error Detection in Tables
Qixu Chen, Yeye He, Raymond Chi-Wing Wong, Weiwei Cui, Song Ge, Haidong Zhang, Dongmei Zhang, Surajit Chaudhuri
摘要
Data cleaning is a long-standing challenge in data management. While powerful logic and statistical algorithms have been developed to detect and repair data errors in tables, existing algorithms predominantly rely on domain-experts to first manually specify data-quality constraints specific to a given table, before data cleaning algorithms can be applied.
In this work, we propose a new class of data-quality constraints that we call Semantic-Domain Constraints, which can be reliably inferred and automatically applied to any tables, without requiring domain-experts to manually specify on a per-table basis. We develop a principled framework to systematically learn such constraints from table corpora using large-scale statistical tests, which can further be distilled into a core set of constraints using our optimization framework, with provable quality guarantees. Extensive evaluations show that this new class of constraints can be used to both (1) directly detect errors on real tables in the wild, and (2) augment existing expert-driven data-cleaning techniques as a new class of complementary constraints.
Our extensively labeled benchmark dataset with 2400 real data columns, as well as our code are available at https://github.com/qixuchen/AutoTest to facilitate future research.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- DeepPrep: An LLM-Powered Agentic System for Autonomous Data PreparationMeihao Fan, Ju Fan, Yuxin Zhang, Shaolei Zhang 等VLDB 2026 · 被引用 4 次
- Table-LLM-Specialist: Language Model Specialists for Tables using Iterative Fine-tuningJunjie Xing, Yeye He, Mengyu Zhou, Haoyu Dong 等EMNLP 2025 · 被引用 2 次
- Auto-Fill: Learning to Predict Missing Values Accurately with Specialist Language ModelsYurong Liu, Yeye He, Haoyu Dong, Junjie Xing 等VLDB 2026
它引用的顶会 Paper12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Can Foundation Models Wrangle Your Data?Avanika Narayan, Ines Chami, Laurel J. Orr, Christopher RéVLDB 2023 · 被引用 325 次
- Annotating Columns with Pre-trained Language ModelsYoshihiko Suhara, Jinfeng Li, Yuliang Li, Dan Zhang 等SIGMOD 2022 · 被引用 81 次
- ArcheType: A Novel Framework for Open-Source Column Type Annotation using Large Language ModelsBenjamin Feuer, Yurong Liu, Chinmay Hegde, Juliana FreireVLDB 2024 · 被引用 33 次
相关 Paper
- How and Why False Denial Constraints are DiscoveredAlbert Martin, Eduardo C. de Almeida, Oscar Romero, Anna QueraltVLDB 2025 · 被引用 1 次
- SCODED: Statistical Constraint Oriented Data Error DetectionJing Nathan Yan, Oliver Schulte, Mohan Zhang, Jiannan Wang 等SIGMOD 2020 · 被引用 32 次
- Discovery of Approximate (and Exact) Denial ConstraintsEduardo H. M. Pena, Eduardo C. de Almeida, Felix NaumannVLDB 2020 · 被引用 79 次
- Guardrail: Automated Integrity Constraint Synthesis From Noisy DataPingchuan Ma, Zhaoyu Wang, Zhenlan Ji, Zongjie Li 等SIGMOD 2026 · 被引用 1 次
- DCDiscover: Mining Threshold Denial Constraints from Time Series DataXiaoou Ding, Muyun Zhou, Yida Liu, Zekai Qian 等ICDE 2025 · 被引用 1 次
