Query-Guided Analysis and Mitigation of Data Verification Errors
Ran Schreiber, Yael Amsterdamer
Abstract
Data verification, the process of labeling data items as correct or incorrect, is a preprocessing step that may critically affect the quality of results in data-driven pipelines. Despite recent advances, verification can still produce erroneous labels that propagate to downstream query results in complex ways. We present a framework that complements existing verification tools by assessing the impact of potential labeling errors on query outputs and guiding additional verification steps to improve result reliability. To this end, we introduce Maximal Error Score (MES), a worst-case uncertainty metric that quantifies the reliability of query output tuples independently of the underlying data distribution. As an auxiliary indicator, we identify risky tuples - input tuples for which reducing label uncertainty may counterintuitively increase the output uncertainty. We then develop efficient algorithms for computing MES and detecting risky tuples, as well as a generic algorithm, named MESReduce, that builds on both indicators and interacts with external verifiers to select effective additional verification steps. We implement our techniques in a prototype system and evaluate them on real and synthetic datasets, demonstrating that MESReduce can substantially and effectively reduce the MES and improve the accuracy of verification results.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f44d3aa3-3d4d-4068-b54f-726bb38f96e8Builds on10
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMsMiao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li et al.ICLR 2024 · 867 citations
- Horizon: Scalable Dependency-driven Data CleaningEl Kindi Rezig, Mourad Ouzzani, Walid G. Aref, Ahmed K. Elmagarmid et al.VLDB 2021 · 95 citations
- Combining Human Predictions with Model Probabilities via Confusion Matrices and CalibrationGavin Kerrigan, Padhraic Smyth, Mark SteyversNeurIPS 2021 · 79 citations
- On Diffusion Modeling for Anomaly DetectionVictor Livernoche, Vineet Jain, Yashar Hezaveh, Siamak RavanbakhshICLR 2024 · 74 citations
- Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language GenerationLorenz Kuhn, Yarin Gal, Sebastian FarquharICLR 2023 · 49 citations
Related papers
- Detecting Misclassification Errors in Neural Networks with a Gaussian Process ModelXin Qiu, Risto MiikkulainenAAAI 2022 · 13 citations
- Conformal Reliability: A New Evaluation Metric for Conditional GenerationYachen Gao, Xinwei Sun, Yikai Wang, Ye Shi et al.ICML 2026
- Better Uncertainty Calibration via Proper Scores for Classification and BeyondSebastian G. Gruber, Florian BuettnerNeurIPS 2022 · 88 citations
- Robust Plan Evaluation based on Approximate Probabilistic Machine LearningAmin Kamali, Verena Kantere, Calisto Zuzarte, Vincent CorvinelliVLDB 2025 · 1 citation
- On Detecting Cherry-picked GeneralizationsYin Lin, Brit Youngmann, Yuval Moskovitch, H. V. Jagadish et al.VLDB 2022 · 18 citations
