DoCoGen: Domain Counterfactual Generation for Low Resource Domain Adaptation
Nitay Calderon, Eyal Ben-David, Amir Feder, Roi Reichart
摘要
Natural language processing (NLP) algorithms have become very successful, but they still struggle when applied to out-of-distribution examples. In this paper we propose a controllable generation approach in order to deal with this domain adaptation (DA) challenge. Given an input text example, our DoCoGen algorithm generates a domain-counterfactual textual example (D-con) - that is similar to the original in all aspects, including the task label, but its domain is changed to a desired one. Importantly, DoCoGen is trained using only unlabeled examples from multiple domains - no NLP task labels or parallel pairs of textual examples and their domain-counterfactuals are required. We show that DoCoGen can generate coherent counterfactuals consisting of multiple sentences. We use the D-cons generated by DoCoGen to augment a sentiment classifier and a multi-label intent classifier in 20 and 78 DA setups, respectively, where source-domain labeled data is scarce. Our model outperforms strong baselines and improves the accuracy of a state-of-the-art unsupervised DA algorithm.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- CEBaB: Estimating the Causal Effects of Real-World Concepts on NLP Model BehaviorEldar David Abraham, Karel D'Oosterlinck, Amir Feder, Yair Ori Gat 等NeurIPS 2022 · 被引用 69 次
- Faithful Explanations of Black-box NLP Models Using LLM-generated CounterfactualsYair Ori Gat, Nitay Calderon, Amir Feder, Alexander Chapanin 等ICLR 2024 · 被引用 55 次
- Counterfactual Generation with Identifiability GuaranteesHanqi Yan, Lingjing Kong, Lin Gui, Yuejie Chi 等NeurIPS 2023 · 被引用 16 次
- Causal-structure Driven Augmentations for Text OOD GeneralizationAmir Feder, Yoav Wald, Claudia Shi, Suchi Saria 等NeurIPS 2023 · 被引用 10 次
- Social Recommendation via Graph-Level Counterfactual AugmentationYinxuan Huang, Ke Liang, Yanyi Huang, Xiang Zeng 等AAAI 2025 · 被引用 8 次
它引用的顶会 Paper9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong 等NeurIPS 2020 · 被引用 2,774 次
- Learning The Difference That Makes A Difference With Counterfactually-Augmented DataDivyansh Kaushik, Eduard H. Hovy, Zachary Chase LiptonICLR 2020 · 被引用 625 次
- Competency Problems: On Finding and Removing Artifacts in Language DataMatt Gardner, William Merrill, Jesse Dodge, Matthew E. Peters 等EMNLP 2021 · 被引用 72 次
- If beam search is the answer, what was the question?Clara Meister, Ryan Cotterell, Tim VieiraEMNLP 2020 · 被引用 26 次
相关 Paper
- Dual Adversarial Co-Learning for Multi-Domain Text ClassificationYuan Wu, Yuhong GuoAAAI 2020 · 被引用 26 次
- Cross-Domain Data Augmentation with Domain-Adaptive Language Modeling for Aspect-Based Sentiment AnalysisJianfei Yu, Qiankun Zhao, Rui XiaACL 2023 · 被引用 19 次
- Prompt-based Distribution Alignment for Domain Generalization in Text ClassificationChen Jia, Yue ZhangEMNLP 2022 · 被引用 4 次
- Deep Domain-Adversarial Image Generation for Domain GeneralisationKaiyang Zhou, Yongxin Yang, Timothy M. Hospedales, Tao XiangAAAI 2020 · 被引用 488 次
- Multi-Source Domain Adaptation for Text Classification via DistanceNet-BanditsHan Guo, Ramakanth Pasunuru, Mohit BansalAAAI 2020 · 被引用 120 次
