Detecting Unassimilated Borrowings in Spanish: An Annotated Corpus and Approaches to Modeling
Elena Álvarez Mellado, Constantine Lignos
摘要
This work presents a new resource for borrowing identification and analyzes the performance and errors of several models on this task. We introduce a new annotated corpus of Spanish newswire rich in unassimilated lexical borrowings—words from one language that are introduced into another without orthographic adaptation—and use it to evaluate how several sequence labeling models (CRF, BiLSTM-CRF, and Transformer-based models) perform. The corpus contains 370,000 tokens and is larger, more borrowing-dense, OOV-rich, and topic-varied than previous corpora available for this task. Our results show that a BiLSTM-CRF model fed with subword embeddings along with either Transformer-based embeddings pretrained on codeswitched data or a combination of contextualized word embeddings outperforms results obtained by a multilingual BERT-based model.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- From English to Code-Switching: Transfer Learning with Strong Morphological CluesGustavo Aguilar, Thamar SolorioACL 2020 · 被引用 1 次
- Exploring morphology-aware tokenization: A case study on Spanish language modelingAlba Táboas García, Piotr Przybyla, Leo WannerEMNLP 2025 · 被引用 1 次
- RoBoCoP: A Comprehensive ROmance BOrrowing COgnate Package and Benchmark for Multilingual Cognate IdentificationLiviu P. Dinu, Ana Sabina Uban, Alina Maria Cristea, Anca P. Dinu 等EMNLP 2023 · 被引用 2 次
- Code-switched inspired losses for spoken dialog representationsPierre Colombo, Emile Chapuis, Matthieu Labeau, Chloé ClavelEMNLP 2021 · 被引用 6 次
- Overlap-based Vocabulary Generation Improves Cross-lingual Transfer Among Related LanguagesVaidehi Patil, Partha P. Talukdar, Sunita SarawagiACL 2022 · 被引用 39 次
