Query-Guided Resolution in Uncertain Databases
Osnat Drien, Matanya Freiman, Antoine Amarilli, Yael Amsterdamer
摘要
We present a novel framework for uncertain data management. We start with a database whose tuple correctness is uncertain and an oracle that can resolve the uncertainty, i.e., decide if a tuple is correct or not. Such an oracle may correspond, e.g., to a data expert or to a crowdsourcing platform. We wish to use the oracle to clean the database with the goal of ensuring the correct answer for specific mission-critical queries. To avoid the prohibitive cost of cleaning the entire database and to minimize the expected number of calls to the oracle, we must carefully select tuples whose resolution would suffice to resolve the uncertainty in query results. In other words, we need a query-guided process for the resolution of uncertain data. We develop an end-to-end solution to this problem, based on the derivation of query answers and on correctness probabilities for the uncertain data. At a high level, we first track Boolean provenance to identify which input tuples contribute to the derivation of each output tuple, and in what ways. We then design an active learning solution for iteratively choosing tuples to resolve, based on the provenance structure and on an evolving estimation of tuple correctness probabilities. We conduct an extensive experimental study to validate our framework in different use cases. P R E P R I N T exhaustive cleaning of the entire database, which can be prohibitively costly, previous work suggested approaches for semi-automated, interactive or crowdpowered cleaning processes (e.g., [8, 11, 19, 20, 70, 22, 32, 60, 64, 81, 93, 96] ). Typically, such processes clean only a part of the data and/or only specific types of data errors, such as constraint violations. In the present work, we study the problem of query-guided uncertainty resolution. We start from a database whose data correctness is uncertain (i.e., which may contain incorrect tuples) and a query (or a set of queries) that capture the data relevant for analysis. Our goal is to identify the precise set of correct query results, by resolving the uncertainty of input tuples. Of course, we can do this by naively verifying every input tuple, thereby achieving a certain and correct input database. However, as mentioned above, verifying all input tuples may be too costly. We thus aim at verifying a subset of the input tuples that suffices to determine the correct results of the given query. As we will show, there exist such subsets that are typically significantly smaller than the full database. Resolving the uncertainty of a tuple is abstractly modeled as a probe to an oracle: in practice, the oracle may be data experts, crowd workers, high-quality external sources, etc. Therefore, our challenge is identifying which tuples to probe in order to minimize the number of oracle calls. Our novel solution addresses this challenge by accounting, in a fine-grained manner, for how the data is derived and for the probabilities of probe answers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper4
- Horizon: Scalable Dependency-driven Data CleaningEl Kindi Rezig, Mourad Ouzzani, Walid G. Aref, Ahmed K. Elmagarmid 等VLDB 2021 · 被引用 95 次
- On Multiple Semantics for Declarative Database RepairsAmir Gilad, Daniel Deutch, Sudeepa RoySIGMOD 2020 · 被引用 20 次
- Enabling Personal Consent in DatabasesGeorge Konstantinidis, Jet Holt, Adriane ChapmanVLDB 2022 · 被引用 17 次
- Compact, Tamper-Resistant Archival of Fine-Grained ProvenanceNan Zheng, Zack IvesVLDB 2021 · 被引用 6 次
相关 Paper
- CrowdRL: An End-to-End Reinforcement Learning Framework for Data LabellingKaiyu Li, Guoliang Li, Yong Wang, Yan Huang 等ICDE 2021 · 被引用 17 次
- Asking the Right Questions to the Right Users: Active Learning with Imperfect OraclesShayok ChakrabortyAAAI 2020 · 被引用 23 次
- Hierarchical Crowdsourcing for Data Labeling with Heterogeneous CrowdHaodi Zhang, Wenxi Huang, Zhenhan Su, Junyang Chen 等ICDE 2023 · 被引用 4 次
- Learn 3D VQA Better with Active Selection and ReannotationShengli Zhou, Yang Liu, Feng ZhengACM MM 2025 · 被引用 2 次
- Putting Things into Context: Rich Explanations for Query Answers using Join GraphsChenjie Li, Zhengjie Miao, Qitian Zeng, Boris Glavic 等SIGMOD 2021 · 被引用 16 次
