Analysis of Candidate Keys in Relational Databases
Zihui Yang, Yuqian Ma, Sebastian Link
Abstract
Discovery algorithms in data profiling return an effective representation of all constraints from a given class, such as uniqueness constraints, that hold on the given dataset. Most of the results, however, do not express meaningful business rules but simply hold incidentally. Incomplete and inconsistent data makes the identification of business rules even harder, leading to approximate algorithms with huge computational complexity and few meaningful constraints among many incidental ones. Research is missing a methodology for analyzing the output of discovery algorithms. In response, we propose a framework for identifying and analyzing candidate keys in relational databases with incomplete or inconsistent data. The approach leverages uniqueness and completeness ratios, threshold-based key filtering, and counter-example analysis to identify meaningful keys, and incorporates specialization techniques and pruning strategies to ensure minimality and efficiency. Experiments demonstrate that key analysis improves F1 measures from state-of-the-art key mining between 36-99%, highlighting the necessity of humans-in-the-loop for decision making. Our pruning strategies yield up to 60% reduction in runtime for key analysis, lowering computational overheads while maintaining result quality.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 2f0cdbbc-1188-47f1-86f9-e344f5dbf8d6Related papers
- Discovering Functional Dependencies through Hitting Set EnumerationTobias Bleifuß, Thorsten Papenbrock, Thomas Bläsius, Martin Schirneck et al.SIGMOD 2024 · 9 citations
- Discovery of Approximate (and Exact) Denial ConstraintsEduardo H. M. Pena, Eduardo C. de Almeida, Felix NaumannVLDB 2020 · 79 citations
- Discovering Approximate Denial Constraints in Large DatabasesAlbert Martin, Eduardo C. de Almeida, Oscar Romero, Anna QueraltVLDB 2026 · 2 citations
- DCDiscover: Mining Threshold Denial Constraints from Time Series DataXiaoou Ding, Muyun Zhou, Yida Liu, Zekai Qian et al.ICDE 2025 · 1 citation
- Fast Algorithms for Denial Constraint DiscoveryEduardo H. M. Pena, Fábio Porto, Felix NaumannVLDB 2023 · 23 citations
