Lune

EMNLP2024Top-tier venue

TEMA: Token Embeddings Mapping for Enriching Low-Resource Language Models

Rodolfo Zevallos, Núria Bel, Mireia Farrús

2024Year

Abstract

The objective of the research we present is to remedy the problem of the low quality of language models for low-resource languages. We introduce an algorithm, the Token Embedding Mapping Algorithm (TEMA), that maps the token embeddings of a richly pre-trained model L1 to a poorly trained model L2, thus creating a richer L2' model. Our experiments show that the L2' model reduces perplexity with respect to the original monolingual model L2, and that for downstream tasks, including SuperGLUE, the results are state-of-the-art or better for the most semantic tasks. The models obtained with TEMA are also competitive or better than multilingual or extended models proposed as solutions for mitigating the low-resource language problems.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 18d5c043-d22b-483a-bffb-fa9d2bea2b53

Builds on10

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines