Lune

ICML2023顶会

Causal Proxy Models for Concept-based Model Explanations

Zhengxuan Wu, Karel D'Oosterlinck, Atticus Geiger, Amir Zur, Christopher Potts

2023年份
40被引次数
7顶会引用

摘要

Explainability methods for NLP systems encounter a version of the fundamental problem of causal inference: for a given ground-truth input text, we never truly observe the counterfactual texts necessary for isolating the causal effects of model representations on outputs. In response, many explainability methods make no use of counterfactual texts, assuming they will be unavailable. In this paper, we show that robust causal explainability methods can be created using approximate counterfactuals, which can be written by humans to approximate a specific counterfactual or simply sampled using metadata-guided heuristics. The core of our proposal is the Causal Proxy Model (CPM). A CPM explains a black-box model N\mathcal{N} because it is trained to have the same actual input/output behavior as N\mathcal{N} while creating neural representations that can be intervened upon to simulate the counterfactual input/output behavior of N\mathcal{N}. Furthermore, we show that the best CPM for N\mathcal{N} performs comparably to N\mathcal{N} in making factual predictions, which means that the CPM can simply replace N\mathcal{N}, leading to more explainable deployed models. Our code is available at https://github.com/frankaging/Causal-Proxy-Model.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper7

问问它们各自怎么用它

它引用的顶会 Paper9

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖