Back-Training excels Self-Training at Unsupervised Domain Adaptation of Question Generation and Passage Retrieval
Devang Kulshreshtha, Robert Belfer, Iulian Vlad Serban, Siva Reddy
摘要
In this work, we introduce back-training, an alternative to self-training for unsupervised domain adaptation (UDA) from source to target domain. While self-training generates synthetic training data where natural inputs are aligned with noisy outputs, back-training results in natural outputs aligned with noisy inputs. This significantly reduces the gap between the target domain and synthetic data distribution, and reduces model overfitting to the source domain. We run UDA experiments on question generation and passage retrieval from the Natural Questions domain to machine learning and biomedical domains. We find that back-training vastly outperforms selftraining by a mean improvement of 7.8 BLEU-4 points on generation, and 17.6% top-20 retrieval accuracy across both domains. We further propose consistency filters to remove low-quality synthetic data before training. We also release a new domain-adaptation dataset-MLQuestions containing 35K unaligned questions, 50K unaligned passages, and 3K aligned question-passage pairs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Response Enhanced Semi-supervised Dialogue Query GenerationJianheng Huang, Ante Wang, Linfeng Gao, Linfeng Song 等AAAI 2024 · 被引用 5 次
- Simulating Bandit Learning from User Feedback for Extractive Question AnsweringGe Gao, Eunsol Choi, Yoav ArtziACL 2022
它引用的顶会 Paper1
相关 Paper
- Unsupervised Data Augmentation with Naive Augmentation and without Unlabeled DataDavid Lowell, Brian E. Howard, Zachary C. Lipton, Byron C. WallaceEMNLP 2021 · 被引用 2 次
- Unsupervised Adaptation of Question Answering Systems via Generative Self-trainingSteven J. Rennie, Etienne Marcheret, Neil Mallinar, David Nahamoo 等EMNLP 2020 · 被引用 11 次
- Focus on Your Target: A Dual Teacher-Student Framework for Domain-adaptive Semantic SegmentationXinyue Huo, Lingxi Xie, Wengang Zhou, Houqiang Li 等ICCV 2023 · 被引用 18 次
- Feature Adaptation of Pre-Trained Language Models across Languages and Domains with Robust Self-TrainingHai Ye, Qingyu Tan, Ruidan He, Juntao Li 等EMNLP 2020 · 被引用 36 次
- SRoUDA: Meta Self-Training for Robust Unsupervised Domain AdaptationWanqing Zhu, Jia-Li Yin, Bo-Hao Chen, Ximeng LiuAAAI 2023 · 被引用 14 次
