Identifying Insufficient Data Coverage in Databases with Multiple Relations
Yin Lin, Yifan Guan, Abolfazl Asudeh, H. V. Jagadish
Abstract
In today's data-driven world, it is critical that we use appropriate datasets for analysis and decision-making. Datasets could be biased because they reflect existing inequalities in the world, due to the data scientists' biased world view, or due to the data scientists' limited control over the data collection process. For these reasons, it is important to ensure adequate data coverage across different groups over the intersection of multiple attributes. Often, the dataset to be analyzed is obtained through complex joins and predicate combinations over multiple relational tables in a database. Due to the sheer data volume we often have to deal with, determining adequate coverage can require an unacceptably long execution time. In this paper, we provide an efficient approach for coverage analysis, given a set of attributes across multiple tables. To identify regions with insufficient coverage in the combinatorially large set of value combinations, we design an index scheme to avoid explicit table joins, achieve efficient memory usage, and support predicate combination at a high level of parallelism. We also propose P-WALK, a priority-based search algorithm, to traverse the lattice space. Since in practice, coverage assessment typically does not require precise COUNT aggregation results, we further present approximate methods to reduce computation time. Experimental evaluation using three real-world datasets shows the effectiveness, efficiency, and accuracy of proposed methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 725a202c-e92e-4ff9-a910-8c3ae7279467Cited by top-tier papers7
- Tailoring Data Source Distributions for Fairness-aware Data IntegrationFatemeh Nargesian, Abolfazl Asudeh, H. V. JagadishVLDB 2021 · 51 citations
- SliceLine: Fast, Linear-Algebra-based Slice Finding for ML Model DebuggingSvetlana Sagadeeva, Matthias BoehmSIGMOD 2021 · 45 citations
- Identifying Insufficient Data Coverage for Ordinal Continuous-Valued AttributesAbolfazl Asudeh, Nima Shahbazi, Zhongjun Jin, H. V. JagadishSIGMOD 2021 · 30 citations
- Fairness-Aware Range Queries for Selecting Unbiased DataSuraj Shetiya, Ian P. Swift, Abolfazl Asudeh, Gautam DasICDE 2022 · 19 citations
- View-based Explanations for Graph Neural NetworksTingyang Chen, Dazhuo Qiu, Yinghui Wu, Arijit Khan et al.SIGMOD 2024 · 17 citations
Related papers
- Efficiently Answering Top-k Window Aggregate Queries: Calculating Coverage Number Sequences over Hierarchical StructuresJianqiu Xu, Raymond Chi-Wing WongICDE 2023
- Bounding and Approximating Intersectional Fairness through Marginal FairnessMathieu Molina, Patrick LoiseauNeurIPS 2022 · 16 citations
- Weighted Set Multi-Cover on Bounded Universe and Applications in Package RecommendationNima Shahbazi, Aryan Esmailpour, Stavros SintosSIGMOD 2026
- Conformal Classification with Equalized Coverage for Adaptively Selected GroupsYanfei Zhou, Matteo SesiaNeurIPS 2024 · 14 citations
- Index Intersection for High-Dimensional Range QueriesMaximilian Berens, Jens TeubnerVLDB 2026
