Improving Constrained Search Results By Data Melioration
Ido Guy, Tova Milo, Slava Novgorodov, Brit Youngmann
摘要
The problem of finding an item-set of maximal aggregated utility that satisfies a set of constraints is at the cornerstone of many search applications. Its classical definition assumes that all the information needed to verify the constraints is explicitly given. However, in real-world databases, the data available on items is often partial. Hence, adequately answering constrained search queries requires the completion of this missing information. A common approach to complete missing data is to employ Machine Learning (ML)-based inference. However, such methods are naturally error-prone. More accurate data can be obtained by asking humans to complete missing information. But, as the number of items in the repository is vast, limiting human effort is crucial. To this end, we introduce the Probabilistic Constrained Search (PCS) problem, which identifies a bounded-size item-set whose data completion is likely to be highly beneficial, as these items are expected to belong to the result set of the constrained search queries in question. We prove PCS to be hard to approximate, and consequently propose a best-effort PTIME heuristic to solve it. We demonstrate the effectiveness and efficiency of our algorithm over real-world datasets and scenarios, showing that our algorithm significantly improves the result sets of constrained search queries, in terms of both utility and constraints satisfaction probability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Guided Exploration of Data SummariesBrit Youngmann, Sihem Amer-Yahia, Aurélien PersonnazVLDB 2022 · 被引用 22 次
- Classifier Construction Under Budget ConstraintsShay Gershtein, Tova Milo, Slava Novgorodov, Kathy RazmadzeSIGMOD 2022 · 被引用 2 次
它引用的顶会 Paper1
相关 Paper
- GoodCore: Data-effective and Data-efficient Machine Learning through Coreset Selection over Incomplete DataChengliang Chai, Jiabin Liu, Nan Tang, Ju Fan 等SIGMOD 2023 · 被引用 37 次
- Contribution Maximization in Probabilistic DatalogTova Milo, Yuval Moskovitch, Brit YoungmannICDE 2020 · 被引用 2 次
- Explaining Missing Data in Graphs: A Constraint-based ApproachQi Song, Peng Lin, Hanchao Ma, Yinghui WuICDE 2021 · 被引用 7 次
- Learning to Learn in Interactive Constraint AcquisitionDimosthenis C. Tsouros, Senne Berden, Tias GunsAAAI 2024 · 被引用 9 次
- Searching a Database of Source Codes Using Contextualized Code SearchRohan Mukherjee, Chris Jermaine, Swarat ChaudhuriVLDB 2020 · 被引用 11 次
