Discovering Approximate Inclusion Dependencies
Qingdong Su, Zhikang Wang, Zijing Tan, Shuai Ma
摘要
Inclusion dependencies (INDs) are widely used in data management tasks. The discovery techniques of INDs have thus received a lot of attention, for discovering INDs valid in data. However, real-world data quality issues may lead to partial violations of INDs. This paper makes the first effort to provide a comprehensive study on the discovery of approximate INDs (AINDs), aiming to identify INDs with error rates below a given threshold. This paper introduces a new definition of AIND based on deletion semantics, in addition to the existing definition based on insertion semantics. A discovery method is developed that can be configured to identify AINDs based on either of these semantics. The method combines partitioning techniques to handle tables that cannot all fit into memory simultaneously, with novel approaches to quantify AIND violations based on partitioned tables. To improve efficiency, the method employs a novel three-layer filtering structure and techniques that can potentially prune invalid candidate AINDs and identify valid AINDs without necessarily processing all tuples. We conduct an extensive experimental evaluation and verify the following: the proposed method significantly outperforms existing methods for AIND discovery based on insertion semantics, the AIND discoveries with insertion and deletion semantics can provide complementary results, and our discovery method can effectively deal with dirty dataset containing various types of errors.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Finding Related Tables in Data Lakes for Interactive Data ScienceYi Zhang, Zachary G. IvesSIGMOD 2020 · 被引用 98 次
- Discovery of Approximate (and Exact) Denial ConstraintsEduardo H. M. Pena, Eduardo C. de Almeida, Felix NaumannVLDB 2020 · 被引用 79 次
- Efficient Joinable Table Discovery in Data Lakes: A High-Dimensional Similarity-Based ApproachYuyang Dong, Kunihiro Takeoka, Chuan Xiao, Masafumi OyamadaICDE 2021 · 被引用 78 次
- Approximate Denial ConstraintsEster Livshits, Alireza Heidari, Ihab F. Ilyas, Benny KimelfeldVLDB 2020 · 被引用 60 次
- Integrating Data Lake TablesAamod Khatiwada, Roee Shraga, Wolfgang Gatterbauer, Renée J. MillerVLDB 2023 · 被引用 59 次
相关 Paper
- Discovering Similarity Inclusion DependenciesYouri Kaminsky, Eduardo H. M. Pena, Felix NaumannSIGMOD 2023 · 被引用 16 次
- Approximate Order Dependency DiscoveryYifeng Jin, Zijing Tan, Weijun Zeng, Shuai MaICDE 2021 · 被引用 7 次
- Fast Incremental Discovery of Pointwise Order DependenciesZijing Tan, Ai Ran, Shuai Ma, Sheng QinVLDB 2020 · 被引用 22 次
- Anytime Algorithms for Approximate Functional DependenciesSanjivni Rana, Junya Ogawa, Suraj Shetiya, Senjuti Basu Roy 等KDD 2025
- DAFDiscover: Robust Mining Algorithm for Dynamic Approximate Functional Dependencies on Dirty DataXiaoou Ding, Yixing Lu, Hongzhi Wang, Chen Wang 等VLDB 2024 · 被引用 4 次
