Back-Training excels Self-Training at Unsupervised Domain Adaptation of Question Generation and Passage Retrieval
Devang Kulshreshtha, Robert Belfer, Iulian Vlad Serban, Siva Reddy
Abstract
In this work, we introduce back-training, an alternative to self-training for unsupervised domain adaptation (UDA) from source to target domain. While self-training generates synthetic training data where natural inputs are aligned with noisy outputs, back-training results in natural outputs aligned with noisy inputs. This significantly reduces the gap between the target domain and synthetic data distribution, and reduces model overfitting to the source domain. We run UDA experiments on question generation and passage retrieval from the Natural Questions domain to machine learning and biomedical domains. We find that back-training vastly outperforms selftraining by a mean improvement of 7.8 BLEU-4 points on generation, and 17.6% top-20 retrieval accuracy across both domains. We further propose consistency filters to remove low-quality synthetic data before training. We also release a new domain-adaptation dataset-MLQuestions containing 35K unaligned questions, 50K unaligned passages, and 3K aligned question-passage pairs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Response Enhanced Semi-supervised Dialogue Query GenerationJianheng Huang, Ante Wang, Linfeng Gao, Linfeng Song et al.AAAI 2024 · 5 citations
- Simulating Bandit Learning from User Feedback for Extractive Question AnsweringGe Gao, Eunsol Choi, Yoav ArtziACL 2022
Builds on1
Related papers
- Unsupervised Data Augmentation with Naive Augmentation and without Unlabeled DataDavid Lowell, Brian E. Howard, Zachary C. Lipton, Byron C. WallaceEMNLP 2021 · 2 citations
- Unsupervised Adaptation of Question Answering Systems via Generative Self-trainingSteven J. Rennie, Etienne Marcheret, Neil Mallinar, David Nahamoo et al.EMNLP 2020 · 11 citations
- Focus on Your Target: A Dual Teacher-Student Framework for Domain-adaptive Semantic SegmentationXinyue Huo, Lingxi Xie, Wengang Zhou, Houqiang Li et al.ICCV 2023 · 18 citations
- Feature Adaptation of Pre-Trained Language Models across Languages and Domains with Robust Self-TrainingHai Ye, Qingyu Tan, Ruidan He, Juntao Li et al.EMNLP 2020 · 36 citations
- SRoUDA: Meta Self-Training for Robust Unsupervised Domain AdaptationWanqing Zhu, Jia-Li Yin, Bo-Hao Chen, Ximeng LiuAAAI 2023 · 14 citations
