The Cost of Representation by Subset Repairs
Yuxi Liu, Fangzhu Shen, Kushagra Ghosh, Amir Gilad, Benny Kimelfeld, Sudeepa Roy
Abstract
Datasets may include errors, and specifically violations of integrity constraints, for various reasons. Standard techniques for "minimalcost" database repairing resolve these violations by aiming for a minimum change in the data, and in the process, may sway representations of different sub-populations. For instance, the repair may end up deleting more females than males, or more tuples from a certain age group or race, due to varying levels of inconsistency in different sub-populations. Such repaired data can mislead consumers when used for analytics, and can lead to biased decisions for downstream machine learning tasks. We study the "cost of representation" in subset repairs for functional dependencies. In simple terms, we target the question of how many additional tuples have to be deleted if we want to satisfy not only the integrity constraints but also representation constraints for given sub-populations. We study the complexity of this problem and compare it with the complexity of optimal subset repairs without representations. While the problem is NP-hard in general, we give polynomial-time algorithms for special cases, and efficient heuristics for general cases. We perform a suite of experiments that show the effectiveness of our algorithms in computing or approximating the cost of representation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6528d97f-76a3-43ee-ab1b-848c70206447Cited by top-tier papers1
Ask how each one uses itBuilds on7
- Retiring Adult: New Datasets for Fair Machine LearningFrances Ding, Moritz Hardt, John Miller, Ludwig SchmidtNeurIPS 2021 · 671 citations
- Automated Feature Engineering for Algorithmic FairnessRicardo Salazar, Felix Neutatz, Ziawasch AbedjanVLDB 2021 · 42 citations
- The Computation of Optimal Subset RepairsDongjing Miao, Zhipeng Cai, Jianzhong Li, Xiangyu Gao et al.VLDB 2020 · 23 citations
- Properties of Inconsistency Measures for DatabasesEster Livshits, Rina Kochirgan, Segev Tsur, Ihab F. Ilyas et al.SIGMOD 2021 · 21 citations
- On Multiple Semantics for Declarative Database RepairsAmir Gilad, Daniel Deutch, Sudeepa RoySIGMOD 2020 · 20 citations
Related papers
- Truth Frequency: Leveraging Dependencies for Subset RepairHaoda Li, Jiahui Chen, Yu Sun, Shaoxu Song et al.ICDE 2026
- Stress-Testing Causal Claims via Cardinality RepairsYarden Gabbay, Haoquan Guan, Shaull Almagor, El Kindi Rezig et al.SIGMOD 2026 · 1 citation
- Representation Matters: Assessing the Importance of Subgroup Allocations in Training DataEsther Rolf, Theodora T. Worledge, Benjamin Recht, Michael I. JordanICML 2021 · 50 citations
- Understanding Fairness and Prediction Error through Subspace Decomposition and Influence AnalysisEnze Shi, Pankaj Bhagwat, Zhixian Yang, Linglong Kong et al.NeurIPS 2025
- Improving Subgroup Robustness via Data SelectionSaachi Jain, Kimia Hamidieh, Kristian Georgiev, Andrew Ilyas et al.NeurIPS 2024 · 17 citations
