Text Counterfactuals via Latent Optimization and Shapley-Guided Search
Xiaoli Z. Fern, Quintin Pope
Abstract
We study the problem of generating counterfactual text for a classifier as a means for understanding and debugging classification. Given a textual input and a classification model, we aim to minimally alter the text to change the model's prediction. White-box approaches have been successfully applied to similar problems in vision where one can directly optimize the continuous input. Optimization-based approaches become difficult in the language domain due to the discrete nature of text. We bypass this issue by directly optimizing in the latent space and leveraging a language model to generate candidate modifications from optimized latent representations. We additionally use Shapley values to estimate the combinatoric effect of multiple changes. We then use these estimates to guide a beam search for the final counterfactual text. We achieve favorable performance compared to recent whitebox and black-box baselines using human and automatic evaluations. Ablation studies show that both latent optimization and the use of Shapley values improve success rate and the quality of the generated counterfactuals.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Towards Model Robustness: Generating Contextual Counterfactuals for Entities in Relation ExtractionMi Zhang, Tieyun Qian, Ting Zhang, Xin MiaoWWW 2023 · 9 citations
- A General Search-Based Framework for Generating Textual Counterfactual ExplanationsDaniel Gilo, Shaul MarkovitchAAAI 2024 · 3 citations
Builds on3
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 1,333 citations
- There and Back Again: Revisiting Backpropagation Saliency MethodsSylvestre-Alvise Rebuffi, Ruth Fong, Xu Ji, Andrea VedaldiCVPR 2020
Related papers
- Generative causal explanations of black-box classifiersMatthew R. O'Shaughnessy, Gregory Canal, Marissa Connor, Christopher Rozell et al.NeurIPS 2020 · 83 citations
- FOCUS: Flexible Optimizable Counterfactual Explanations for Tree EnsemblesAna Lucic, Harrie Oosterhuis, Hinda Haned, Maarten de RijkeAAAI 2022 · 87 citations
- CF-OPT: Counterfactual Explanations for Structured PredictionGermain Vivier-Ardisson, Alexandre Forel, Axel Parmentier, Thibaut VidalICML 2024 · 3 citations
- Ordered Counterfactual Explanation by Mixed-Integer Linear OptimizationKentaro Kanamori, Takuya Takagi, Ken Kobayashi, Yuichi Ike et al.AAAI 2021 · 135 citations
- Trustworthy Actionable PerturbationsJesse Friedbaum, Sudarshan Adiga, Ravi TandonICML 2024 · 3 citations
