Learning The Difference That Makes A Difference With Counterfactually-Augmented Data
Divyansh Kaushik, Eduard H. Hovy, Zachary Chase Lipton
Abstract
Despite alarm over the reliance of machine learning systems on so-called spurious patterns, the term lacks coherent meaning in standard statistical frameworks. However, the language of causality offers clarity: spurious associations are due to confounding (e.g., a common cause), but not direct or indirect causal effects. In this paper, we focus on natural language processing, introducing methods and resources for training models less sensitive to spurious patterns. Given documents and their initial labels, we task humans with revising each document so that it (i) accords with a counterfactual target label; (ii) retains internal coherence; and (iii) avoids unnecessary changes. Interestingly, on sentiment analysis and natural language inference tasks, classifiers trained on original data fail on their counterfactually-revised counterparts and vice versa. Classifiers trained on combined datasets perform remarkably well, just shy of those specialized to either domain. While classifiers trained on either original or manipulated data alone are sensitive to spurious features (e.g., mentions of genre), models trained on the combined data are less sensitive to this signal. Both datasets are publicly available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4b4fffa2-93c1-4004-947b-e21a855259f0Cited by top-tier papers170
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- The Hateful Memes Challenge: Detecting Hate Speech in Multimodal MemesDouwe Kiela, Hamed Firooz, Aravind Mohan, Vedanuj Goswami et al.NeurIPS 2020 · 1,022 citations
- Aligning AI With Shared Human ValuesDan Hendrycks, Collin Burns, Steven Basart, Andrew Critch et al.ICLR 2021 · 878 citations
- Investigating Gender Bias in Language Models Using Causal Mediation AnalysisJesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian et al.NeurIPS 2020 · 851 citations
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal et al.ACL 2020 · 602 citations
Related papers
- Counterfactual Invariance to Spurious Correlations in Text ClassificationVictor Veitch, Alexander D'Amour, Steve Yadlowsky, Jacob EisensteinNeurIPS 2021 · 108 citations
- Robustness to Spurious Correlations in Text Classification via Automatically Generated CounterfactualsZhao Wang, Aron CulottaAAAI 2021 · 114 citations
- Towards Robust Classification Model by Counterfactual and Invariant Data GenerationChun-Hao Chang, George-Alexandru Adam, Anna GoldenbergCVPR 2021
- Explaining the Efficacy of Counterfactually Augmented DataDivyansh Kaushik, Amrith Setlur, Eduard H. Hovy, Zachary Chase LiptonICLR 2021 · 89 citations
- C2L: Causally Contrastive Learning for Robust Text ClassificationSeungtaek Choi, Myeongho Jeong, Hojae Han, Seung-won HwangAAAI 2022 · 52 citations
