Neural Mask Generator: Learning to Generate Adaptive Word Maskings for Language Model Adaptation
Minki Kang, Moonsu Han, Sung Ju Hwang
摘要
We propose a method to automatically generate a domain-and task-adaptive maskings of the given text for self-supervised pre-training, such that we can effectively adapt the language model to a particular target task (e.g. question answering). Specifically, we present a novel reinforcement learning-based framework which learns the masking policy, such that using the generated masks for further pre-training of the target language model helps improve task performance on unseen texts. We use off-policy actor-critic with entropy regularization and experience replay for reinforcement learning, and propose a Transformer-based policy network that can consider the relative importance of words in a given text. We validate our Neural Mask Generator (NMG) on several question answering and text classification datasets using BERT and DistilBERT as the language models, on which it outperforms rule-based masking strategies, by automatically learning optimal adaptive maskings. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- When FLUE Meets FLANG: Benchmarks and Large Pretrained Language Model for Financial DomainRaj Sanjay Shah, Kunal Chawla, Dheeraj Eidnani, Agam Shah 等EMNLP 2022 · 被引用 63 次
- Meta-learning to Improve Pre-trainingAniruddh Raghu, Jonathan Lorraine, Simon Kornblith, Matthew McDermott 等NeurIPS 2021 · 被引用 39 次
- Self-Distillation for Further Pre-training of TransformersSeanie Lee, Minki Kang, Juho Lee, Sung Ju Hwang 等ICLR 2023 · 被引用 4 次
- On the Influence of Masking Policies in Intermediate Pre-trainingQinyuan Ye, Belinda Z. Li, Sinong Wang, Benjamin Bolte 等EMNLP 2021 · 被引用 2 次
它引用的顶会 Paper5
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
- Data Valuation using Reinforcement LearningJinsung Yoon, Sercan Ömer Arik, Tomas PfisterICML 2020 · 被引用 236 次
- Don't Stop Pretraining: Adapt Language Models to Domains and TasksSuchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo 等ACL 2020 · 被引用 93 次
- Span Selection Pre-training for Question AnsweringMichael R. Glass, Alfio Gliozzo, Rishav Chakravarti, Anthony Ferritto 等ACL 2020 · 被引用 9 次
相关 Paper
- PALM: Pre-training an Autoencoding&Autoregressive Language Model for Context-conditioned GenerationBin Bi, Chenliang Li, Chen Wu, Ming Yan 等EMNLP 2020 · 被引用 41 次
- Exploiting Structured Knowledge in Text via Graph-Guided Representation LearningTao Shen, Yi Mao, Pengcheng He, Guodong Long 等EMNLP 2020 · 被引用 60 次
- Effective Unsupervised Domain Adaptation with Adversarially Trained Language ModelsThuy-Trang Vu, Dinh Phung, Gholamreza HaffariEMNLP 2020 · 被引用 19 次
- Adapting a Language Model While Preserving its General KnowledgeZixuan Ke, Yijia Shao, Haowei Lin, Hu Xu 等EMNLP 2022 · 被引用 6 次
- UniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-TrainingHangbo Bao, Li Dong, Furu Wei, Wenhui Wang 等ICML 2020 · 被引用 423 次
