Polyjuice: Generating Counterfactuals for Explaining, Evaluating, and Improving Models
Tongshuang Wu, Marco Túlio Ribeiro, Jeffrey Heer, Daniel S. Weld
Abstract
While counterfactual examples are useful for analysis and training of NLP models, current generation methods either rely on manual labor to create very few counterfactuals, or only instantiate limited types of perturbations such as paraphrases or word substitutions. We present Polyjuice, a general-purpose counterfactual generator that allows for control over perturbation types and locations, trained by finetuning GPT-2 on multiple datasets of paired sentences. We show that Polyjuice produces diverse sets of realistic counterfactuals, which in turn are useful in various distinct applications: improving training and evaluation on three different tasks (with around 70% less annotation effort than manual generation), augmenting state-of-the-art explanation techniques, and supporting systematic counterfactual error analysis by revealing behaviors easily missed by human experts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5e6b4804-bbf2-4db1-b790-18421c2984b7Cited by top-tier papers64
- AI Chains: Transparent and Controllable Human-AI Interaction by Chaining Large Language Model PromptsTongshuang Wu, Michael Terry, Carrie Jun CaiCHI 2022 · 465 citations
- Red Teaming Language Models with Language ModelsEthan Perez, Saffron Huang, H. Francis Song, Trevor Cai et al.EMNLP 2022 · 239 citations
- Counterfactual Invariance to Spurious Correlations in Text ClassificationVictor Veitch, Alexander D'Amour, Steve Yadlowsky, Jacob EisensteinNeurIPS 2021 · 108 citations
- Adaptive Testing and Debugging of NLP ModelsMarco Túlio Ribeiro, Scott M. LundbergACL 2022 · 99 citations
- Tailor: Generating and Perturbing Text with Semantic ControlsAlexis Ross, Tongshuang Wu, Hao Peng, Matthew E. Peters et al.ACL 2022 · 85 citations
Builds on10
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- Plug and Play Language Models: A Simple Approach to Controlled Text GenerationSumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung et al.ICLR 2020 · 1,166 citations
- Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team PerformanceGagan Bansal, Tongshuang Wu, Joyce Zhou, Raymond Fok et al.CHI 2021 · 713 citations
- Learning The Difference That Makes A Difference With Counterfactually-Augmented DataDivyansh Kaushik, Eduard H. Hovy, Zachary Chase LiptonICLR 2020 · 625 citations
- Evaluating Explainable AI: Which Algorithmic Explanations Help Users Predict Model Behavior?Peter Hase, Mohit BansalACL 2020 · 216 citations
Related papers
- People Make Better Edits: Measuring the Efficacy of LLM-Generated Counterfactually Augmented Data for Harmful Language DetectionIndira Sen, Dennis Assenmacher, Mattia Samory, Isabelle Augenstein et al.EMNLP 2023 · 7 citations
- JoPA: Explaining Large Language Model's Generation via Joint Prompt AttributionYurui Chang, Bochuan Cao, Yujia Wang, Jinghui Chen et al.ACL 2025
- Exploring the Efficacy of Automatically Generated Counterfactuals for Sentiment AnalysisLinyi Yang, Jiazheng Li, Padraig Cunningham, Yue Zhang et al.ACL 2021
- DISCO: Distilling Counterfactuals with Large Language ModelsZeming Chen, Qiyue Gao, Antoine Bosselut, Ashish Sabharwal et al.ACL 2023 · 27 citations
- Counterfactual Generator: A Weakly-Supervised Method for Named Entity RecognitionXiangji Zeng, Yunliang Li, Yuchen Zhai, Yin ZhangEMNLP 2020 · 55 citations
