Taming Pre-trained Language Models with N-gram Representations for Low-Resource Domain Adaptation
Shizhe Diao, Ruijia Xu, Hongjin Su, Yilei Jiang, Yan Song, Tong Zhang
摘要
Large pre-trained models such as BERT are known to improve different downstream NLP tasks, even when such a model is trained on a generic domain. Moreover, recent studies have shown that when large domain-specific corpora are available, continued pre-training on domain-specific data can further improve the performance of in-domain tasks. However, this practice requires significant domainspecific data and computational resources which may not always be available. In this paper, we aim to adapt a generic pretrained model with a relatively small amount of domain-specific data. We demonstrate that by explicitly incorporating the multi-granularity information of unseen and domain-specific words via the adaptation of (word based) ngrams, the performance of a generic pretrained model can be greatly improved. Specifically, we introduce a Transformer-based Domainaware N-gram Adaptor, T-DNA, to effectively learn and incorporate the semantic representation of different combinations of words in the new domain. Experimental results illustrate the effectiveness of T-DNA on eight lowresource downstream tasks from four domains. We show that T-DNA is able to achieve significant improvements compared to existing methods on most tasks using limited data with lower computational costs. Moreover, further analyses demonstrate the importance and effectiveness of both unseen words and the information of different granularities. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- LISA: Layerwise Importance Sampling for Memory-Efficient Large Language Model Fine-TuningRui Pan, Xiang Liu, Shizhe Diao, Renjie Pi 等NeurIPS 2024 · 被引用 124 次
- Sparse Invariant Risk MinimizationXiao Zhou, Yong Lin, Weizhong Zhang, Tong ZhangICML 2022 · 被引用 85 次
- Model Agnostic Sample Reweighting for Out-of-Distribution LearningXiao Zhou, Yong Lin, Renjie Pi, Weizhong Zhang 等ICML 2022 · 被引用 73 次
- ORGAN: Observation-Guided Radiology Report Generation via Tree ReasoningWenjun Hou, Kaishuai Xu, Yi Cheng, Wenjie Li 等ACL 2023 · 被引用 36 次
- Mixture-of-Domain-Adapters: Decoupling and Injecting Domain Knowledge to Pre-trained Language Models' MemoriesShizhe Diao, Tianyang Xu, Ruijia Xu, Jiawei Wang 等ACL 2023 · 被引用 17 次
它引用的顶会 Paper6
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Generalization through Memorization: Nearest Neighbor Language ModelsUrvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer 等ICLR 2020 · 被引用 1,038 次
- Improving Chinese Word Segmentation with Wordhood Memory NetworksYuanhe Tian, Yan Song, Fei Xia, Tong Zhang 等ACL 2020 · 被引用 95 次
- Don't Stop Pretraining: Adapt Language Models to Domains and TasksSuchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo 等ACL 2020 · 被引用 93 次
相关 Paper
- Entity Extraction in Low Resource Domains with Selective Pre-training of Large Language ModelsAniruddha Mahapatra, Sharmila Reddy Nangi, Aparna Garimella, Anandhavelu NatarajanEMNLP 2022 · 被引用 4 次
- Self-Distillation for Further Pre-training of TransformersSeanie Lee, Minki Kang, Juho Lee, Sung Ju Hwang 等ICLR 2023 · 被引用 4 次
- KnowDA: All-in-One Knowledge Mixture Model for Data Augmentation in Low-Resource NLPYufei Wang, Jiayi Zheng, Can Xu, Xiubo Geng 等ICLR 2023 · 被引用 2 次
- MetaMT, a Meta Learning Method Leveraging Multiple Domain Data for Low Resource Machine TranslationRumeng Li, Xun Wang, Hong YuAAAI 2020 · 被引用 42 次
- Meta-Transfer Learning for Low-Resource Abstractive SummarizationYi-Syuan Chen, Hong-Han ShuaiAAAI 2021 · 被引用 41 次
