DoCoGen: Domain Counterfactual Generation for Low Resource Domain Adaptation
Nitay Calderon, Eyal Ben-David, Amir Feder, Roi Reichart
Abstract
Natural language processing (NLP) algorithms have become very successful, but they still struggle when applied to out-of-distribution examples. In this paper we propose a controllable generation approach in order to deal with this domain adaptation (DA) challenge. Given an input text example, our DoCoGen algorithm generates a domain-counterfactual textual example (D-con) - that is similar to the original in all aspects, including the task label, but its domain is changed to a desired one. Importantly, DoCoGen is trained using only unlabeled examples from multiple domains - no NLP task labels or parallel pairs of textual examples and their domain-counterfactuals are required. We show that DoCoGen can generate coherent counterfactuals consisting of multiple sentences. We use the D-cons generated by DoCoGen to augment a sentiment classifier and a multi-label intent classifier in 20 and 78 DA setups, respectively, where source-domain labeled data is scarce. Our model outperforms strong baselines and improves the accuracy of a state-of-the-art unsupervised DA algorithm.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers13
- CEBaB: Estimating the Causal Effects of Real-World Concepts on NLP Model BehaviorEldar David Abraham, Karel D'Oosterlinck, Amir Feder, Yair Ori Gat et al.NeurIPS 2022 · 69 citations
- Faithful Explanations of Black-box NLP Models Using LLM-generated CounterfactualsYair Ori Gat, Nitay Calderon, Amir Feder, Alexander Chapanin et al.ICLR 2024 · 55 citations
- Counterfactual Generation with Identifiability GuaranteesHanqi Yan, Lingjing Kong, Lin Gui, Yuejie Chi et al.NeurIPS 2023 · 16 citations
- Causal-structure Driven Augmentations for Text OOD GeneralizationAmir Feder, Yoav Wald, Claudia Shi, Suchi Saria et al.NeurIPS 2023 · 10 citations
- Social Recommendation via Graph-Level Counterfactual AugmentationYinxuan Huang, Ke Liang, Yanyi Huang, Xiang Zeng et al.AAAI 2025 · 8 citations
Builds on9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong et al.NeurIPS 2020 · 2,774 citations
- Learning The Difference That Makes A Difference With Counterfactually-Augmented DataDivyansh Kaushik, Eduard H. Hovy, Zachary Chase LiptonICLR 2020 · 625 citations
- Competency Problems: On Finding and Removing Artifacts in Language DataMatt Gardner, William Merrill, Jesse Dodge, Matthew E. Peters et al.EMNLP 2021 · 72 citations
- If beam search is the answer, what was the question?Clara Meister, Ryan Cotterell, Tim VieiraEMNLP 2020 · 26 citations
Related papers
- Dual Adversarial Co-Learning for Multi-Domain Text ClassificationYuan Wu, Yuhong GuoAAAI 2020 · 26 citations
- Cross-Domain Data Augmentation with Domain-Adaptive Language Modeling for Aspect-Based Sentiment AnalysisJianfei Yu, Qiankun Zhao, Rui XiaACL 2023 · 19 citations
- Prompt-based Distribution Alignment for Domain Generalization in Text ClassificationChen Jia, Yue ZhangEMNLP 2022 · 4 citations
- Deep Domain-Adversarial Image Generation for Domain GeneralisationKaiyang Zhou, Yongxin Yang, Timothy M. Hospedales, Tao XiangAAAI 2020 · 488 citations
- Multi-Source Domain Adaptation for Text Classification via DistanceNet-BanditsHan Guo, Ramakanth Pasunuru, Mohit BansalAAAI 2020 · 120 citations
