Effective Unsupervised Domain Adaptation with Adversarially Trained Language Models
Thuy-Trang Vu, Dinh Phung, Gholamreza Haffari
摘要
Recent work has shown the importance of adaptation of broad-coverage contextualised embedding models on the domain of the target task of interest. Current self-supervised adaptation methods are simplistic, as the training signal comes from a small percentage of randomly masked-out tokens. In this paper, we show that careful masking strategies can bridge the knowledge gap of masked language models (MLMs) about the domains more effectively by allocating self-supervision where it is needed. Furthermore, we propose an effective training strategy by adversarially masking out those tokens which are harder to reconstruct by the underlying MLM. The adversarial objective leads to a challenging combinatorial optimisation problem over subsets of tokens, which we tackle efficiently through relaxation to a variational lower-bound and dynamic programming. On six unsupervised domain adaptation tasks involving named entity recognition, our method strongly outperforms the random masking strategy and achieves up to +1.64 F1 score improvements.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Active Learning on Pre-trained Language Model with Task-Independent Triplet LossSeungmin Seo, Donghyun Kim, Youbin Ahn, Kyong-Ho LeeAAAI 2022 · 被引用 21 次
- Curriculum CycleGAN for Textual Sentiment Domain Adaptation with Multiple SourcesSicheng Zhao, Yang Xiao, Jiang Guo, Xiangyu Yue 等WWW 2021 · 被引用 19 次
- Out-of-Distribution Generalization in Natural Language Processing: Past, Present, and FutureLinyi Yang, Yaoxian Song, Xuan Ren, Chenyang Lyu 等EMNLP 2023 · 被引用 12 次
- Adapt in Contexts: Retrieval-Augmented Domain Adaptation via In-Context LearningQuanyu Long, Wenya Wang, Sinno Jialin PanEMNLP 2023 · 被引用 11 次
- Generalised Unsupervised Domain Adaptation of Neural Machine Translation with Cross-Lingual Data SelectionThuy-Trang Vu, Xuanli He, Dinh Q. Phung, Gholamreza HaffariEMNLP 2021 · 被引用 2 次
它引用的顶会 Paper2
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
- Don't Stop Pretraining: Adapt Language Models to Domains and TasksSuchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo 等ACL 2020 · 被引用 93 次
相关 Paper
- Exploiting Structured Knowledge in Text via Graph-Guided Representation LearningTao Shen, Yi Mao, Pengcheng He, Guodong Long 等EMNLP 2020 · 被引用 60 次
- Neural Mask Generator: Learning to Generate Adaptive Word Maskings for Language Model AdaptationMinki Kang, Moonsu Han, Sung Ju HwangEMNLP 2020 · 被引用 12 次
- InforMask: Unsupervised Informative Masking for Language Model PretrainingNafis Sadeq, Canwen Xu, Julian J. McAuleyEMNLP 2022 · 被引用 12 次
- Mask-Align: Self-Supervised Neural Word AlignmentChi Chen, Maosong Sun, Yang LiuACL 2021
- Learning Dynamic Contextualised Word Embeddings via Template-based Temporal AdaptationXiaohang Tang, Yi Zhou, Danushka BollegalaACL 2023 · 被引用 2 次
