Improving Pretraining Techniques for Code-Switched NLP
Richeek Das, Sahasra Ranjan, Shreya Pathak, Preethi Jyothi
摘要
Pretrained models are a mainstay in modern NLP applications. Pretraining requires access to large volumes of unlabeled text. While monolingual text is readily available for many of the world’s languages, access to large quantities of code-switched text (i.e., text with tokens of multiple languages interspersed within a sentence) is much more scarce. Given this resource constraint, the question of how pretraining using limited amounts of code-switched text could be altered to improve performance for code-switched NLP becomes important to tackle. In this paper, we explore different masked language modeling (MLM) pretraining techniques for code-switched text that are cognizant of language boundaries prior to masking. The language identity of the tokens can either come from human annotators, trained language classifiers, or simple relative frequency-based estimates. We also present an MLM variant by introducing a residual connection from an earlier layer in the pretrained model that uniformly boosts performance on downstream tasks. Experiments on two downstream tasks, Question Answering (QA) and Sentiment Analysis (SA), involving four code-switched language pairs (Hindi-English, Spanish-English, Tamil-English, Malayalam-English) yield relative improvements of up to 5.8 and 2.7 F1 scores on QA (Hindi-English) and SA (Tamil-English), respectively, compared to standard pretraining techniques. To understand our task improvements better, we use a series of probes to study what additional information is encoded by our pretraining techniques and also introduce an auxiliary loss function that explicitly models language identification to further aid the residual MLM variants.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper6
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- Improving Massively Multilingual Neural Machine Translation and Zero-Shot TranslationBiao Zhang, Philip Williams, Ivan Titov, Rico SennrichACL 2020 · 被引用 213 次
- GLUECoS: An Evaluation Benchmark for Code-Switched NLPSimran Khanuja, Sandipan Dandapat, Anirudh Srinivasan, Sunayana Sitaram 等ACL 2020 · 被引用 12 次
- Learning Better Masking for Better Language Model Pre-trainingDongjie Yang, Zhuosheng Zhang, Hai ZhaoACL 2023 · 被引用 9 次
相关 Paper
- Multilingual Large Language Models Are Not (Yet) Code-SwitchersRuochen Zhang, Samuel Cahyawijaya, Jan Christian Blaise Cruz, Genta Indra Winata 等EMNLP 2023 · 被引用 19 次
- Alternating Language Modeling for Cross-Lingual Pre-TrainingJian Yang, Shuming Ma, Dongdong Zhang, Shuangzhi Wu 等AAAI 2020 · 被引用 94 次
- From Machine Translation to Code-Switching: Generating High-Quality Code-Switched TextIshan Tarunesh, Syamantak Kumar, Preethi JyothiACL 2021
- From English to Code-Switching: Transfer Learning with Strong Morphological CluesGustavo Aguilar, Thamar SolorioACL 2020 · 被引用 1 次
- CSP: Code-Switching Pre-training for Neural Machine TranslationZhen Yang, Bojie Hu, Ambyera Han, Shen Huang 等EMNLP 2020 · 被引用 62 次
