Conformance Constraint Discovery: Measuring Trust in Data-Driven Systems
Anna Fariha, Ashish Tiwari, Arjun Radhakrishna, Sumit Gulwani, Alexandra Meliou
Abstract
The reliability of inferences made by data-driven systems hinges on the data's continued conformance to the systems' initial settings and assumptions. When serving data (on which we want to apply inference) deviates from the profile of the initial training data, the outcome of inference becomes unreliable. We introduce conformance constraints, a new data profiling primitive tailored towards quantifying the degree of non-conformance, which can effectively characterize if inference over that tuple is untrustworthy. Conformance constraints are constraints over certain arithmetic expressions (called projections) involving the numerical attributes of a dataset, which existing data profiling primitives such as functional dependencies and denial constraints cannot model.
The key finding we present is that projections that incur low variance on a dataset construct effective conformance constraints. This principle yields the surprising result that low-variance components of a principal component analysis, which are usually discarded for dimensionality reduction, generate stronger conformance constraints than the high-variance components. Based on this result, we provide a highly scalable and efficient technique-linear in data size and cubic in the number of attributes-for discovering conformance constraints for a dataset. To measure the degree of a tuple's non-conformance with respect to a dataset, we propose a quantitative semantics that captures how much a tuple violates the conformance constraints of that dataset. We demonstrate the value of conformance constraints on two applications: trusted machine learning and data drift. We empirically show that conformance constraints offer mechanisms to (1) reliably detect tuples on which the inference of a machine-learned model should not be trusted, and (2) quantify data drift more accurately than the state of the art.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 468bced9-1764-43e6-9fb6-60a9e51423b1Cited by top-tier papers7
- Exathlon: A Benchmark for Explainable Anomaly Detection over Time SeriesVincent Jacob, Fei Song, Arnaud Stiegler, Bijan Rad et al.VLDB 2021 · 97 citations
- MTSClean: Efficient Constraint-based Cleaning for Multi-Dimensional Time Series DataXiaoou Ding, Yichen Song, Hongzhi Wang, Chen Wang et al.VLDB 2024 · 9 citations
- DataPrism: Exposing Disconnect between Data and SystemsSainyam Galhotra, Anna Fariha, Raoni Lourenço, Juliana Freire et al.SIGMOD 2022 · 9 citations
- Data-Semantics-Aware Recommendation of Diverse Pivot TablesWhanhee Cho, Anna FarihaSIGMOD 2026 · 4 citations
- Inferring Data Preconditions from Deep Learning Models for Trustworthy Prediction in DeploymentShibbir Ahmed, Hongyang Gao, Hridesh RajanICSE 2024 · 3 citations
Builds on5
- Discovery of Approximate (and Exact) Denial ConstraintsEduardo H. M. Pena, Eduardo C. de Almeida, Felix NaumannVLDB 2020 · 79 citations
- Approximate Denial ConstraintsEster Livshits, Alireza Heidari, Ihab F. Ilyas, Benny KimelfeldVLDB 2020 · 60 citations
- A Statistical Perspective on Discovering Functional Dependencies in Noisy DataYunjia Zhang, Zhihan Guo, Theodoros RekatsinasSIGMOD 2020 · 45 citations
- Pattern Functional Dependencies for Data CleaningAbdulhakim Ali Qahtan, Nan Tang, Mourad Ouzzani, Yang Cao et al.VLDB 2020 · 42 citations
- SCODED: Statistical Constraint Oriented Data Error DetectionJing Nathan Yan, Oliver Schulte, Mohan Zhang, Jiannan Wang et al.SIGMOD 2020 · 32 citations
Related papers
- Data-SUITE: Data-centric identification of in-distribution incongruous examplesNabeel Seedat, Jonathan Crabbé, Mihaela van der SchaarICML 2022 · 16 citations
- Non-Invasive Fairness in Learning Through the Lens of Data DriftKe Yang, Alexandra MeliouICDE 2024 · 2 citations
- Conformal Mixed-Integer Constraint Learning with Feasibility GuaranteesDaniel Ovalle, Lorenz T. Biegler, Ignacio E. Grossmann, Carl D. Laird et al.NeurIPS 2025 · 2 citations
- Discovering Approximate Denial Constraints in Large DatabasesAlbert Martin, Eduardo C. de Almeida, Oscar Romero, Anna QueraltVLDB 2026 · 2 citations
- How and Why False Denial Constraints are DiscoveredAlbert Martin, Eduardo C. de Almeida, Oscar Romero, Anna QueraltVLDB 2025 · 1 citation
