Can We Improve Model Robustness through Secondary Attribute Counterfactuals?
Ananth Balashankar, Xuezhi Wang, Ben Packer, Nithum Thain, Ed H. Chi, Alex Beutel
Abstract
Developing robust NLP models that perform well on many, even small, slices of data is a significant but important challenge, with implications from fairness to general reliability. To this end, recent research has explored how models rely on spurious correlations, and how counterfactual data augmentation (CDA) can mitigate such issues. In this paper we study how and why modeling counterfactuals over multiple attributes can go significantly further in improving model performance. We propose RDI, a context-aware methodology which takes into account the impact of secondary attributes on the model's predictions and increases sensitivity for secondary attributes over reweighted counterfactually augmented data. By implementing RDI in the context of toxicity detection, we find that accounting for secondary attributes can significantly improve robustness, with improvements in sliced accuracy on the original dataset up to 7% compared to existing robustness methods. We also demonstrate that RDI generalizes to the coreference resolution task and provide guidelines to extend this to other tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cefcd0c6-6264-4442-bf44-b3874bae13dbCited by top-tier papers4
- High Fidelity Image Counterfactuals with Probabilistic Causal ModelsFabio De Sousa Ribeiro, Tian Xia, Miguel Monteiro, Nick Pawlowski et al.ICML 2023 · 68 citations
- Probing Classifiers are Unreliable for Concept Removal and DetectionAbhinav Kumar, Chenhao Tan, Amit SharmaNeurIPS 2022 · 46 citations
- Doubly Abductive Counterfactual Inference for Text-Based Image EditingXue Song, Jiequan Cui, Hanwang Zhang, Jingjing Chen et al.CVPR 2024 · 7 citations
- Looking at Radiology Report Generation through a Causal Lens: A SurveySatyam Kumar, Kaustubh Shivshankar Shejole, Pushpak BhattacharyyaACL 2026
Builds on7
- Climbing towards NLU: On Meaning, Form, and Understanding in the Age of DataEmily M. Bender, Alexander KollerACL 2020 · 914 citations
- Learning The Difference That Makes A Difference With Counterfactually-Augmented DataDivyansh Kaushik, Eduard H. Hovy, Zachary Chase LiptonICLR 2020 · 625 citations
- Language (Technology) is Power: A Critical Survey of "Bias" in NLPSu Lin Blodgett, Solon Barocas, Hal Daumé III, Hanna M. WallachACL 2020 · 68 citations
- Demographics Should Not Be the Reason of Toxicity: Mitigating Discrimination in Text Classifications with Instance WeightingGuanhua Zhang, Bing Bai, Junqi Zhang, Kun Bai et al.ACL 2020 · 56 citations
- Beyond Accuracy: Behavioral Testing of NLP Models with CheckListMarco Túlio Ribeiro, Tongshuang Wu, Carlos Guestrin, Sameer SinghACL 2020 · 51 citations
Related papers
- C2L: Causally Contrastive Learning for Robust Text ClassificationSeungtaek Choi, Myeongho Jeong, Hojae Han, Seung-won HwangAAAI 2022 · 52 citations
- How Does Counterfactually Augmented Data Impact Models for Social Computing Constructs?Indira Sen, Mattia Samory, Fabian Flöck, Claudia Wagner et al.EMNLP 2021 · 20 citations
- Towards Model Robustness: Generating Contextual Counterfactuals for Entities in Relation ExtractionMi Zhang, Tieyun Qian, Ting Zhang, Xin MiaoWWW 2023 · 9 citations
- Explaining the Efficacy of Counterfactually Augmented DataDivyansh Kaushik, Amrith Setlur, Eduard H. Hovy, Zachary Chase LiptonICLR 2021 · 89 citations
- Exploring the Efficacy of Automatically Generated Counterfactuals for Sentiment AnalysisLinyi Yang, Jiazheng Li, Padraig Cunningham, Yue Zhang et al.ACL 2021
