Effective Unsupervised Domain Adaptation with Adversarially Trained Language Models
Thuy-Trang Vu, Dinh Phung, Gholamreza Haffari
Abstract
Recent work has shown the importance of adaptation of broad-coverage contextualised embedding models on the domain of the target task of interest. Current self-supervised adaptation methods are simplistic, as the training signal comes from a small percentage of randomly masked-out tokens. In this paper, we show that careful masking strategies can bridge the knowledge gap of masked language models (MLMs) about the domains more effectively by allocating self-supervision where it is needed. Furthermore, we propose an effective training strategy by adversarially masking out those tokens which are harder to reconstruct by the underlying MLM. The adversarial objective leads to a challenging combinatorial optimisation problem over subsets of tokens, which we tackle efficiently through relaxation to a variational lower-bound and dynamic programming. On six unsupervised domain adaptation tasks involving named entity recognition, our method strongly outperforms the random masking strategy and achieves up to +1.64 F1 score improvements.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Active Learning on Pre-trained Language Model with Task-Independent Triplet LossSeungmin Seo, Donghyun Kim, Youbin Ahn, Kyong-Ho LeeAAAI 2022 · 21 citations
- Curriculum CycleGAN for Textual Sentiment Domain Adaptation with Multiple SourcesSicheng Zhao, Yang Xiao, Jiang Guo, Xiangyu Yue et al.WWW 2021 · 19 citations
- Out-of-Distribution Generalization in Natural Language Processing: Past, Present, and FutureLinyi Yang, Yaoxian Song, Xuan Ren, Chenyang Lyu et al.EMNLP 2023 · 12 citations
- Adapt in Contexts: Retrieval-Augmented Domain Adaptation via In-Context LearningQuanyu Long, Wenya Wang, Sinno Jialin PanEMNLP 2023 · 11 citations
- Generalised Unsupervised Domain Adaptation of Neural Machine Translation with Cross-Lingual Data SelectionThuy-Trang Vu, Xuanli He, Dinh Q. Phung, Gholamreza HaffariEMNLP 2021 · 2 citations
Builds on2
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- Don't Stop Pretraining: Adapt Language Models to Domains and TasksSuchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo et al.ACL 2020 · 93 citations
Related papers
- Exploiting Structured Knowledge in Text via Graph-Guided Representation LearningTao Shen, Yi Mao, Pengcheng He, Guodong Long et al.EMNLP 2020 · 60 citations
- Neural Mask Generator: Learning to Generate Adaptive Word Maskings for Language Model AdaptationMinki Kang, Moonsu Han, Sung Ju HwangEMNLP 2020 · 12 citations
- InforMask: Unsupervised Informative Masking for Language Model PretrainingNafis Sadeq, Canwen Xu, Julian J. McAuleyEMNLP 2022 · 12 citations
- Mask-Align: Self-Supervised Neural Word AlignmentChi Chen, Maosong Sun, Yang LiuACL 2021
- Learning Dynamic Contextualised Word Embeddings via Template-based Temporal AdaptationXiaohang Tang, Yi Zhou, Danushka BollegalaACL 2023 · 2 citations
