DILBERT: Customized Pre-Training for Domain Adaptation with Category Shift, with an Application to Aspect Extraction
Entony Lekhtman, Yftah Ziser, Roi Reichart
摘要
The rise of pre-trained language models has yielded substantial progress in the vast majority of Natural Language Processing (NLP) tasks. However, a generic approach towards the pre-training procedure can naturally be sub-optimal in some cases. Particularly, finetuning a pre-trained language model on a source domain and then applying it to a different target domain, results in a sharp performance decline of the eventual classifier for many source-target domain pairs. Moreover, in some NLP tasks, the output categories substantially differ between domains, making adaptation even more challenging. This, for example, happens in the task of aspect extraction, where the aspects of interest of reviews of, e.g., restaurants or electronic devices may be very different. This paper presents a new fine-tuning scheme for BERT, which aims to address the above challenges. We name this scheme DILBERT: Domain Invariant Learning with BERT, and customize it for aspect extraction in the unsupervised domain adaptation setting. DILBERT harnesses the categorical information of both the source and the target domains to guide the pre-training process towards a more domain and category invariant representation, thus closing the gap between the domains. We show that DILBERT yields substantial improvements over state-ofthe-art baselines while using a fraction of the unlabeled data, particularly in more challenging domain adaptation setups. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Cross-Domain Data Augmentation with Domain-Adaptive Language Modeling for Aspect-Based Sentiment AnalysisJianfei Yu, Qiankun Zhao, Rui XiaACL 2023 · 被引用 19 次
- Improving Temporal Generalization of Pre-trained Language Models with Lexical Semantic ChangeZhaochen Su, Zecheng Tang, Xinyan Guan, Lijun Wu 等EMNLP 2022 · 被引用 11 次
- Cooperative and Adversarial Learning: Co-enhancing Discriminability and Transferability in Domain AdaptationHui Sun, Zheng Xie, Xin-Ye Li, Ming LiAAAI 2023 · 被引用 5 次
- DoCoGen: Domain Counterfactual Generation for Low Resource Domain AdaptationNitay Calderon, Eyal Ben-David, Amir Feder, Roi ReichartACL 2022
- Predicting Text Preference Via Structured Comparative ReasoningJing Nathan Yan, Tianqi Liu, Justin T. Chiu, Jiaming Shen 等ACL 2024
它引用的顶会 Paper1
相关 Paper
- Adversarial and Domain-Aware BERT for Cross-Domain Sentiment AnalysisChunning Du, Haifeng Sun, Jingyu Wang, Qi Qi 等ACL 2020 · 被引用 165 次
- Meta Fine-Tuning Neural Language Models for Multi-Domain Text MiningChengyu Wang, Minghui Qiu, Jun Huang, Xiaofeng HeEMNLP 2020 · 被引用 19 次
- Domain-oriented Language Modeling with Adaptive Hybrid Masking and Optimal Transport AlignmentDenghui Zhang, Zixuan Yuan, Yanchi Liu, Hao Liu 等KDD 2021 · 被引用 7 次
- Adapting a Language Model While Preserving its General KnowledgeZixuan Ke, Yijia Shao, Haowei Lin, Hu Xu 等EMNLP 2022 · 被引用 6 次
- Feature Adaptation of Pre-Trained Language Models across Languages and Domains with Robust Self-TrainingHai Ye, Qingyu Tan, Ruidan He, Juntao Li 等EMNLP 2020 · 被引用 36 次
