Evaluating Bivariate Causal Statements Based on Mutual Compatibility
Erik Jahn, Dominik Janzing
摘要
For many real-world systems, causal ground truth is difficult to obtain, making claims about causal effects hard to assess. We develop methods for evaluating collections of bivariate causal statements, one for each pair of variables in a fixed system. In the setting of acyclic linear statements, any such collection can be extended to a unique multivariate causal model, but we argue that this induced model is implausible if it imposes substantial additional confounding to explain observed correlations. We introduce a compatibility score that quantifies this notion of plausibility, notably without relying on the faithfulness assumption. Additionally, we define an incompatibility score for purely graphical bivariate causal statements, based on global consistency constraints that are derived from acyclicity and faithfulness assumptions. We give theoretical and empirical evidence that both scores can successfully distinguish correct from incorrect causal statements in generic settings. Moreover, we demonstrate the practical applicability of our methods by analyzing causal claims made by large language models. Our work aims to provide a foundation for assessing the reliability of causal information derived from human experts or artificial intelligence in settings where alternative forms of validation are unavailable.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- Independent mechanism analysis, a new concept?Luigi Gresele, Julius von Kügelgen, Vincent Stimper, Bernhard Schölkopf 等NeurIPS 2021 · 被引用 133 次
- Toward Falsifying Causal Graphs Using a Permutation-Based TestElias Eulig, Atalanti-Anastasia Mastakouri, Patrick Blöbaum, Michaela Hardt 等AAAI 2025 · 被引用 21 次
- Detecting and Measuring Confounding Using Causal Mechanism ShiftsAbbavaram Gowtham Reddy, Vineeth N. BalasubramanianNeurIPS 2024 · 被引用 7 次
相关 Paper
- Causal Order: The Key to Leveraging Imperfect Experts in Causal InferenceAniket Vashishtha, Abbavaram Gowtham Reddy, Abhinav Kumar, Saketh Bachu 等ICLR 2025
- Walk the Talk? Measuring the Faithfulness of Large Language Model ExplanationsKatie Matton, Robert Osazuwa Ness, John V. Guttag, Emre KicimanICLR 2025
- Inference-Time Conformal Reasoning with Valid Factuality Control for Large Language ModelsTing Wang, Yuanjie Shi, Yan Yan, Huan ZhangICML 2026
- Compositional Causal Reasoning Evaluation in Language ModelsJacqueline R. M. A. Maasch, Alihan Hüyük, Xinnuo Xu, Aditya V. Nori 等ICML 2025
- A Causal Lens for Evaluating Faithfulness MetricsKerem Zaman, Shashank SrivastavaEMNLP 2025
