GLUECoS: An Evaluation Benchmark for Code-Switched NLP
Simran Khanuja, Sandipan Dandapat, Anirudh Srinivasan, Sunayana Sitaram, Monojit Choudhury
Abstract
Code-switching is the use of more than one language in the same conversation or utterance. Recently, multilingual contextual embedding models, trained on multiple monolingual corpora, have shown promising results on cross-lingual and multilingual tasks. We present an evaluation benchmark, GLUECoS, for code-switched languages, that spans several NLP tasks in English-Hindi and English-Spanish. Specifically, our evaluation benchmark includes Language Identification from text, POS tagging, Named Entity Recognition, Sentiment Analysis, Question Answering and a new task for code-switching, Natural Language Inference. We present results on all these tasks using cross-lingual word embedding models and multilingual models. In addition, we fine-tune multilingual models on artificially generated code-switched data. Although multilingual models perform significantly better than cross-lingual models, our results show that in most tasks, across both language pairs, multilingual models fine-tuned on code-switched data perform best, showing that multilingual models can be further optimized for code-switching tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 65b763c4-41bd-4a61-9556-18e4caca241aCited by top-tier papers16
- IndoNLG: Benchmark and Resources for Evaluating Indonesian Natural Language GenerationSamuel Cahyawijaya, Genta Indra Winata, Bryan Wilie, Karissa Vincentio et al.EMNLP 2021 · 85 citations
- How Linguistically Fair Are Multilingual Pre-Trained Language Models?Monojit Choudhury, Amit DeshpandeAAAI 2021 · 59 citations
- Few-shot Controllable Style Transfer for Low-Resource Multilingual SettingsKalpesh Krishna, Deepak Nathani, Xavier Garcia, Bidisha Samanta et al.ACL 2022 · 28 citations
- Multilingual Large Language Models Are Not (Yet) Code-SwitchersRuochen Zhang, Samuel Cahyawijaya, Jan Christian Blaise Cruz, Genta Indra Winata et al.EMNLP 2023 · 19 citations
- Toward Language Justice: Exploring Multilingual Captioning for AccessibilityAashaka Desai, Rahaf Alharbi, Stacy Hsueh, Richard E. Ladner et al.CHI 2025 · 11 citations
Related papers
- CoCoa: An Encoder-Decoder Model for Controllable Code-switched GenerationSneha Mondal, Ritika, Shreya Pathak, Preethi Jyothi et al.EMNLP 2022 · 5 citations
- Improving Pretraining Techniques for Code-Switched NLPRicheek Das, Sahasra Ranjan, Shreya Pathak, Preethi JyothiACL 2023 · 3 citations
- From English to Code-Switching: Transfer Learning with Strong Morphological CluesGustavo Aguilar, Thamar SolorioACL 2020 · 1 citation
- From Machine Translation to Code-Switching: Generating High-Quality Code-Switched TextIshan Tarunesh, Syamantak Kumar, Preethi JyothiACL 2021
- GupShup: Summarizing Open-Domain Code-Switched ConversationsLaiba Mehnaz, Debanjan Mahata, Rakesh Gosangi, Uma Sushmitha Gunturi et al.EMNLP 2021 · 13 citations
