Polyjuice: Generating Counterfactuals for Explaining, Evaluating, and Improving Models
Tongshuang Wu, Marco Túlio Ribeiro, Jeffrey Heer, Daniel S. Weld
摘要
While counterfactual examples are useful for analysis and training of NLP models, current generation methods either rely on manual labor to create very few counterfactuals, or only instantiate limited types of perturbations such as paraphrases or word substitutions. We present Polyjuice, a general-purpose counterfactual generator that allows for control over perturbation types and locations, trained by finetuning GPT-2 on multiple datasets of paired sentences. We show that Polyjuice produces diverse sets of realistic counterfactuals, which in turn are useful in various distinct applications: improving training and evaluation on three different tasks (with around 70% less annotation effort than manual generation), augmenting state-of-the-art explanation techniques, and supporting systematic counterfactual error analysis by revealing behaviors easily missed by human experts.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper64
- AI Chains: Transparent and Controllable Human-AI Interaction by Chaining Large Language Model PromptsTongshuang Wu, Michael Terry, Carrie Jun CaiCHI 2022 · 被引用 465 次
- Red Teaming Language Models with Language ModelsEthan Perez, Saffron Huang, H. Francis Song, Trevor Cai 等EMNLP 2022 · 被引用 239 次
- Counterfactual Invariance to Spurious Correlations in Text ClassificationVictor Veitch, Alexander D'Amour, Steve Yadlowsky, Jacob EisensteinNeurIPS 2021 · 被引用 108 次
- Adaptive Testing and Debugging of NLP ModelsMarco Túlio Ribeiro, Scott M. LundbergACL 2022 · 被引用 99 次
- Tailor: Generating and Perturbing Text with Semantic ControlsAlexis Ross, Tongshuang Wu, Hao Peng, Matthew E. Peters 等ACL 2022 · 被引用 85 次
它引用的顶会 Paper10
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- Plug and Play Language Models: A Simple Approach to Controlled Text GenerationSumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung 等ICLR 2020 · 被引用 1,166 次
- Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team PerformanceGagan Bansal, Tongshuang Wu, Joyce Zhou, Raymond Fok 等CHI 2021 · 被引用 713 次
- Learning The Difference That Makes A Difference With Counterfactually-Augmented DataDivyansh Kaushik, Eduard H. Hovy, Zachary Chase LiptonICLR 2020 · 被引用 625 次
- Evaluating Explainable AI: Which Algorithmic Explanations Help Users Predict Model Behavior?Peter Hase, Mohit BansalACL 2020 · 被引用 216 次
相关 Paper
- People Make Better Edits: Measuring the Efficacy of LLM-Generated Counterfactually Augmented Data for Harmful Language DetectionIndira Sen, Dennis Assenmacher, Mattia Samory, Isabelle Augenstein 等EMNLP 2023 · 被引用 7 次
- JoPA: Explaining Large Language Model's Generation via Joint Prompt AttributionYurui Chang, Bochuan Cao, Yujia Wang, Jinghui Chen 等ACL 2025
- Exploring the Efficacy of Automatically Generated Counterfactuals for Sentiment AnalysisLinyi Yang, Jiazheng Li, Padraig Cunningham, Yue Zhang 等ACL 2021
- DISCO: Distilling Counterfactuals with Large Language ModelsZeming Chen, Qiyue Gao, Antoine Bosselut, Ashish Sabharwal 等ACL 2023 · 被引用 27 次
- Counterfactual Generator: A Weakly-Supervised Method for Named Entity RecognitionXiangji Zeng, Yunliang Li, Yuchen Zhai, Yin ZhangEMNLP 2020 · 被引用 55 次
