Pattern Functional Dependencies for Data Cleaning
Abdulhakim Ali Qahtan, Nan Tang, Mourad Ouzzani, Yang Cao, Michael Stonebraker
摘要
Patterns (or regex-based expressions) are widely used to constrain the format of a domain (or a column), e.g., a Year column should contain only four digits, and thus a value like "1980-" might be a typo. Moreover, integrity constraints (ICs) defined over multiple columns, such as (conditional) functional dependencies and denial constraints, e.g., a ZIP code uniquely determines a city in the UK, have been widely used in data cleaning. However, a promising, but not yet explored, direction is to combine regex- and IC-based theories to capture data dependencies involving partial attribute values. For example, in an employee ID such as"F-9-107", "F" is sufficient to determine the finance department. Inspired by the above observation, we propose a novel class of ICs, called pattern functional dependencies (PFDs), to model fine-grained data dependencies gleaned from partial attribute values. These dependencies cannot be modeled using traditional ICs, such as (conditional) functional dependencies, which work on entire attribute values. We also present a set of axioms for the inference of PFDs, analogous to Armstrong's axioms for FDs, and study the complexity of consistency and implication analysis of PFDs. Moreover, we devise an effective algorithm to automatically discover PFDs even in the presence of errors in the data. Our extensive experiments on 15 real-world datasets show that our approach can effectively discover valid and useful PFDs over dirty data, which can then be used to detect data errors that are hard to capture by other types of ICs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- RPT: Relational Pre-trained Transformer Is Almost All You Need towards Democratizing Data PreparationNan Tang, Ju Fan, Fangyi Li, Jianhong Tu 等VLDB 2021 · 被引用 92 次
- Self-supervised and Interpretable Data Cleaning with Sequence Generative Adversarial NetworksJinfeng Peng, Derong Shen, Nan Tang, Tieying Liu 等VLDB 2023 · 被引用 24 次
- Parallel Rule Discovery from Large Datasets by SamplingWenfei Fan, Ziyan Han, Yaoshu Wang, Min XieSIGMOD 2022 · 被引用 21 次
- Conformance Constraint Discovery: Measuring Trust in Data-Driven SystemsAnna Fariha, Ashish Tiwari, Arjun Radhakrishna, Sumit Gulwani 等SIGMOD 2021 · 被引用 18 次
- Efficient Validation of SHACL Shapes with ReasoningJin Ke, Zenon G. Zacouris, Maribel AcostaVLDB 2024 · 被引用 9 次
相关 Paper
- Repairing Entities using Star Constraints in Multirelational GraphsPeng Lin, Qi Song, Yinghui Wu, Jiaxing PiICDE 2020 · 被引用 7 次
- Horizon: Scalable Dependency-driven Data CleaningEl Kindi Rezig, Mourad Ouzzani, Walid G. Aref, Ahmed K. Elmagarmid 等VLDB 2021 · 被引用 95 次
- Inconsistency Detection with Temporal Graph Functional DependenciesMorteza Alipour Langouri, Adam Mansfield, Fei Chiang, Yinghui WuICDE 2023 · 被引用 4 次
- TSDDISCOVER: Discovering Data Dependency for Time Series DataXiaoou Ding, Yingze Li, Hongzhi Wang, Chen Wang 等ICDE 2024 · 被引用 11 次
- Approximate Denial ConstraintsEster Livshits, Alireza Heidari, Ihab F. Ilyas, Benny KimelfeldVLDB 2020 · 被引用 60 次
