Counterfactual Invariance to Spurious Correlations in Text Classification
Victor Veitch, Alexander D'Amour, Steve Yadlowsky, Jacob Eisenstein
Abstract
Informally, a 'spurious correlation' is the dependence of a model on some aspect of the input data that an analyst thinks shouldn't matter. In machine learning, these have a know-it-when-you-see-it character; e.g., changing the gender of a sentence's subject changes a sentiment predictor's output. To check for spurious correlations, we can 'stress test' models by perturbing irrelevant parts of input data and seeing if model predictions change. In this paper, we study stress testing using the tools of causal inference. We introduce counterfactual invariance as a formalization of the requirement that changing irrelevant parts of the input shouldn't change model predictions. We connect counterfactual invariance to out-of-domain model performance, and provide practical schemes for learning (approximately) counterfactual invariant predictors (without access to counterfactual examples). It turns out that both the means and implications of counterfactual invariance depend fundamentally on the true underlying causal structure of the data-in particular, whether the label causes the features or the features cause the label. Distinct causal structures require distinct regularization schemes to induce counterfactual invariance. Similarly, counterfactual invariance implies different domain shift guarantees depending on the underlying causal structure. This theory is supported by empirical results on text classification.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e845f256-0344-4b85-a314-f2d6d62d3abcCited by top-tier papers28
- Probing Classifiers are Unreliable for Concept Removal and DetectionAbhinav Kumar, Chenhao Tan, Amit SharmaNeurIPS 2022 · 46 citations
- Causal Modelling Agents: Causal Graph Discovery through Synergising Metadata- and Data-driven ReasoningAhmed Abdulaal, Adamos Hadjivasiliou, Nina Montaña Brown, Tiantian He et al.ICLR 2024 · 43 citations
- When Does Group Invariant Learning Survive Spurious Correlations?Yimeng Chen, Ruibin Xiong, Zhi-Ming Ma, Yanyan LanNeurIPS 2022 · 30 citations
- Moderately Distributional Exploration for Domain GeneralizationRui Dai, Yonggang Zhang, Zhen Fang, Bo Han et al.ICML 2023 · 28 citations
- Causally motivated multi-shortcut identification and removalJiayun Zheng, Maggie MakarNeurIPS 2022 · 27 citations
Builds on8
- Learning The Difference That Makes A Difference With Counterfactually-Augmented DataDivyansh Kaushik, Eduard H. Hovy, Zachary Chase LiptonICLR 2020 · 625 citations
- Weakly-Supervised Disentanglement Without CompromisesFrancesco Locatello, Ben Poole, Gunnar Rätsch, Bernhard Schölkopf et al.ICML 2020 · 361 citations
- Representation Learning via Invariant Causal MechanismsJovana Mitrovic, Brian McWilliams, Jacob C. Walker, Lars Holger Buesing et al.ICLR 2021 · 281 citations
- Model-Based Domain GeneralizationAlexander Robey, George J. Pappas, Hamed HassaniNeurIPS 2021 · 167 citations
- Counterfactuals uncover the modular structure of deep generative modelsMichel Besserve, Arash Mehrjou, Rémy Sun, Bernhard SchölkopfICLR 2020 · 109 citations
Related papers
- Towards Robust Classification Model by Counterfactual and Invariant Data GenerationChun-Hao Chang, George-Alexandru Adam, Anna GoldenbergCVPR 2021
- Are All Spurious Features in Natural Language Alike? An Analysis through a Causal LensNitish Joshi, Xiang Pan, He HeEMNLP 2022 · 19 citations
- Counterfactual Reasoning for Out-of-distribution Multimodal Sentiment AnalysisTeng Sun, Wenjie Wang, Liqiang Jing, Yiran Cui et al.ACM MM 2022 · 65 citations
- Recovering Latent Causal Factor for Generalization to Distributional ShiftsXinwei Sun, Botong Wu, Xiangyu Zheng, Chang Liu et al.NeurIPS 2021 · 73 citations
- Robustness to Spurious Correlations in Text Classification via Automatically Generated CounterfactualsZhao Wang, Aron CulottaAAAI 2021 · 114 citations
