False Friends or Cognates? A Cross-lingual Semantic Ambiguity Evaluation for Galician, Portuguese and Spanish
Marta Vázquez Abuín, José Camacho-Collados, Marcos García
摘要
The linguistic proximity between Galician, Portuguese, and Spanish results in a lexical overlap that often conceals semantic interference. This is particularly evident in false friends, posing a challenge for NLP systems. In this work, we assess whether state-of-the-art language models can identify and process false friends among these languages. We introduce six cross-lingual datasets -created manually or using semi-automatic methods, with all instances being carefully verified-covering cognates and false friends. We evaluate a broad range of encoder and decoder models of varying sizes via zero-shot and few-shot settings. Our results highlight the challenging nature of the task, but also show the clear progress made by LLMs in recent years, particularly those of a larger size, with smaller language models struggling on the task. Notably, unlike other tasks where language distance poses additional challenges, we find that linguistic proximity itself introduces errors: closely related language pairs tend to perform worse, reflecting the challenge of semantic discrimination due to lexical overlap.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- Don't Trust ChatGPT when your Question is not in English: A Study of Multilingual Abilities and Types of LLMsXiang Zhang, Senyu Li, Bradley Hauer, Ning Shi 等EMNLP 2023 · 被引用 58 次
- XL-WiC: A Multilingual Benchmark for Evaluating Semantic ContextualizationAlessandro Raganato, Tommaso Pasini, José Camacho-Collados, Mohammad Taher PilehvarEMNLP 2020 · 被引用 2 次
- Friend or Foe? A Computational Investigation of Semantic False Friends across Romance LanguagesAna Sabina Uban, Liviu P. Dinu, Ioan-Bogdan Iordache, Simona Georgescu 等EMNLP 2025
相关 Paper
- Exploring the Representation of Word Meanings in Context: A Case Study on Homonymy and SynonymyMarcos GarcíaACL 2021
- FOOL ME IF YOU CAN! An Adversarial Dataset to Investigate the Robustness of LMs in Word Sense DisambiguationMohamad Ballout, Anne Dedert, Nohayr Abdelmoneim, Ulf Krumnack 等EMNLP 2024 · 被引用 2 次
- Cross-Lingual Pitfalls: Automatic Probing Cross-Lingual Weakness of Multilingual Large Language ModelsZixiang Xu, Yanbo Wang, Yue Huang, Xiuying Chen 等ACL 2025 · 被引用 5 次
- RoBoCoP: A Comprehensive ROmance BOrrowing COgnate Package and Benchmark for Multilingual Cognate IdentificationLiviu P. Dinu, Ana Sabina Uban, Alina Maria Cristea, Anca P. Dinu 等EMNLP 2023 · 被引用 2 次
- Multilingual Transfer Learning for QA using Translation as Data AugmentationMihaela A. Bornea, Lin Pan, Sara Rosenthal, Radu Florian 等AAAI 2021 · 被引用 45 次
