Latxa: An Open Language Model and Evaluation Suite for Basque
Julen Etxaniz, Oscar Sainz, Naiara Miguel, Itziar Aldabe, German Rigau, Eneko Agirre, Aitor Ormazabal, Mikel Artetxe, Aitor Soroa
Abstract
We introduce Latxa, a family of large language models for Basque ranging from 7 to 70 billion parameters. Latxa is based on Llama 2, which we continue pretraining on a new Basque corpus comprising 4.3M documents and 4.2B tokens. Addressing the scarcity of high-quality benchmarks for Basque, we further introduce 4 multiple choice evaluation datasets: EusProficiency, comprising 5,169 questions from official language proficiency exams; EusReading, comprising 352 reading comprehension questions; EusTrivia, comprising 1,715 trivia questions from 5 knowledge areas; and EusExams, comprising 16,774 questions from public examinations. In our extensive evaluation, Latxa outperforms all previous open models we compare to by a large margin. In addition, it is competitive with GPT-4 Turbo in language proficiency and understanding, despite lagging behind in reading comprehension and knowledgeintensive tasks. Both the Latxa family of models, as well as our new pretraining corpora and evaluation datasets, are publicly available under open licenses. 1 Our suite enables reproducible research on methods to build LLMs for lowresource languages. * Equal contribution. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6712bc0e-142d-446d-bd7d-16f85e24a90fCited by top-tier papers5
- LLäMmlein: Transparent, Compact and Competitive German-Only Language Models from ScratchJan Pfister, Julia Wunderle, Andreas HothoACL 2025 · 7 citations
- Instructing Large Language Models for Low-Resource Languages: A Systematic Study for BasqueOscar Sainz, Naiara Pérez, Julen Etxaniz, Joseba Fernandez de Landa et al.EMNLP 2025
- ScholaWrite: A Dataset of End-to-End Scholarly WritingKhanh Chi Le, Linghe Wang, Minhwa Lee, Ross Volkov et al.ACL 2026
- INCLUDE: Evaluating Multilingual Language Understanding with Regional KnowledgeAngelika Romanou, Negar Foroutan, Anna Sotnikova, Zeming Chen et al.ICLR 2025
- K-HALU: Multiple Answer Korean Hallucination Benchmark for Large Language ModelsJaehyung Seo, Heuiseok LimICLR 2025
Builds on9
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- Few-shot Learning with Multilingual Generative Language ModelsXi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang et al.EMNLP 2022 · 113 citations
- Translation Artifacts in Cross-lingual Transfer LearningMikel Artetxe, Gorka Labaka, Eneko AgirreEMNLP 2020 · 68 citations
- The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language VariantsLucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe et al.ACL 2024 · 30 citations
Related papers
- Truth Knows No Language: Evaluating Truthfulness Beyond EnglishBlanca Calvo Figueras, Eneko Sagarzazu, Julen Etxaniz, Jeremy Barnes et al.ACL 2025
- La Leaderboard: A Large Language Model Leaderboard for Spanish Varieties and Languages of Spain and Latin AmericaMaría Grandury, Javier Aula-Blasco, Júlia Falcão, Clémentine Fourrier et al.ACL 2025
- Sinhala Encoder-only Language Models and EvaluationTharindu Ranasinghe, Hansi Hettiarachchi, Nadeesha Chathurangi Naradde Vidana Pathirana, Damith Premasiri et al.ACL 2025 · 5 citations
- Multi-LMentry: Can Multilingual LLMs Solve Elementary Tasks Across Languages?Luca Moroni, Javier Aula-Blasco, Simone Conia, Irene Baucells et al.EMNLP 2025
- TUMLU: A Unified and Native Language Understanding Benchmark for Turkic LanguagesJafar Isbarov, Arofat Akhundjanova, Mammad Hajili, Kavsar Huseynova et al.ACL 2025
