Stress-Testing Causal Claims via Cardinality Repairs
Yarden Gabbay, Haoquan Guan, Shaull Almagor, El Kindi Rezig, Brit Youngmann, Babak Salimi
Abstract
Causal analyses derived from observational data underpin high-stakes decisions in domains such as healthcare, public policy, and economics. Yet such conclusions can be surprisingly fragile: even minor data errors - duplicate records, or entry mistakes - may drastically alter causal relationships. This raises a fundamental question: how robust is a causal claim to small, targeted modifications in the data? Addressing this question is essential for ensuring the reliability, interpretability, and reproducibility of empirical findings. We introduce SubCure , a framework for robustness auditing via cardinality repairs . Given a causal query and a user-specified target range for the estimated effect, SubCure identifies a small set of tuples or subpopulations whose removal shifts the estimate into the desired range. This process not only quantifies the sensitivity of causal conclusions but also pinpoints the specific regions of the data that drive those conclusions. We formalize this problem under both tuple- and pattern-level deletion settings and show both are NP-complete. To scale to large datasets, we develop efficient algorithms that incorporate machine unlearning techniques to incrementally update causal estimates without retraining from scratch. We evaluate SubCure across four real-world datasets covering diverse application domains. In each case, it uncovers compact, high-impact subsets whose removal significantly shifts the causal conclusions—revealing vulnerabilities that traditional methods fail to detect. Our results demonstrate that cardinality repair is a powerful and general-purpose tool for stress-testing causal analyses and guarding against misleading claims rooted in ordinary data imperfections.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e534c1f4-b19d-4544-9c8d-bd03b936a17eBuilds on18
- Certified Data Removal from Machine Learning ModelsChuan Guo, Tom Goldstein, Awni Y. Hannun, Laurens van der MaatenICML 2020 · 633 citations
- DeltaGrad: Rapid retraining of machine learning modelsYinjun Wu, Edgar Dobriban, Susan B. DavidsonICML 2020 · 262 citations
- If Influence Functions are the Answer, Then What is the Question?Juhan Bae, Nathan Ng, Alston Lo, Marzyeh Ghassemi et al.NeurIPS 2022 · 185 citations
- Horizon: Scalable Dependency-driven Data CleaningEl Kindi Rezig, Mourad Ouzzani, Walid G. Aref, Ahmed K. Elmagarmid et al.VLDB 2021 · 95 citations
- Explaining Black-Box Algorithms Using Probabilistic Contrastive CounterfactualsSainyam Galhotra, Romila Pradhan, Babak SalimiSIGMOD 2021 · 85 citations
Related papers
- Stress-Testing ML Pipelines with Adversarial Data CorruptionJiongli Zhu, Geyang Xu, Felipe Lorenzi, Boris Glavic et al.VLDB 2025 · 2 citations
- The Cost of Representation by Subset RepairsYuxi Liu, Fangzhu Shen, Kushagra Ghosh, Amir Gilad et al.VLDB 2025 · 4 citations
- Truth Frequency: Leveraging Dependencies for Subset RepairHaoda Li, Jiahui Chen, Yu Sun, Shaoxu Song et al.ICDE 2026
- Forgetting by Pruning: Data Deletion in Join Cardinality EstimationChaowei He, Yuanjun Liu, Qingzhi Ma, Shenyuan Ren et al.AAAI 2026
- DataPrism: Exposing Disconnect between Data and SystemsSainyam Galhotra, Anna Fariha, Raoni Lourenço, Juliana Freire et al.SIGMOD 2022 · 9 citations
