SkMTEB: Slovak Massive Text Embedding Benchmark and Model Adaptation
Marek Suppa, Andrej Ridzik, Daniel Hládek, Natália Knazeková, Viktoria Ondrejova
Abstract
We introduce SkMTEB, the first comprehensive MTEB-style text embedding benchmark for Slovak, a low-resource West Slavic language, comprising 31 datasets across 7 task types -- nearly 4 the depth of existing multilingual benchmark coverage for Slovak. Our evaluation of 31 embedding models reveals that large instruction-tuned multilingual models achieve the strongest performance, while existing Slovak-specific models trained for NLU tasks transfer poorly to embedding tasks. To address the need for efficient, locally-deployable Slovak embeddings, we develop e5-sk-small (45M parameters) and e5-sk-large (365M) by applying vocabulary trimming and fine-tuning to Multilingual E5 models. Despite size reductions of up to 62%, our open-source models achieve competitive performance with proprietary APIs while remaining locally deployable for semantic search and retrieval-augmented generation (RAG). We release the benchmark, models, datasets, and code openly, hoping our approach offers a replicable path for other under-resourced languages.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on4
- Improving Massively Multilingual Neural Machine Translation and Zero-Shot TranslationBiao Zhang, Philip Williams, Ivan Titov, Rico SennrichACL 2020 · 213 citations
- Making Monolingual Sentence Embeddings Multilingual using Knowledge DistillationNils Reimers, Iryna GurevychEMNLP 2020 · 54 citations
- The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language VariantsLucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe et al.ACL 2024 · 30 citations
- Language-agnostic BERT Sentence EmbeddingFangxiaoyu Feng, Yinfei Yang, Daniel Cer, Naveen Arivazhagan et al.ACL 2022
Related papers
- MMTEB: Massive Multilingual Text Embedding BenchmarkKenneth C. Enevoldsen, Isaac Chung, Imene Kerboua, Márton Kardos et al.ICLR 2025 · 10 citations
- Generative Representational Instruction TuningNiklas Muennighoff, Hongjin Su, Liang Wang, Nan Yang et al.ICLR 2025
- Are We on the Right Way to Assess Document Retrieval-Augmented Generation?Wenxuan Shen, Mingjia Wang, Yaochen Wang, Dongping Chen et al.AAAI 2026
- Mieb: Massive Image Embedding BenchmarkChenghao Xiao, Isaac Chung, Imene Kerboua, Jamie Stirling et al.ICCV 2025
- MAIR: A Massive Benchmark for Evaluating Instructed RetrievalWeiwei Sun, Zhengliang Shi, Wu Long, Lingyong Yan et al.EMNLP 2024 · 1 citation
