Lune

EMNLP2025Top-tier venue

TounsiBench: Benchmarking Large Language Models for Tunisian Arabic

Souha Hassine, Asma Arrak, Marouene Addhoum, Steven R. Wilson

2025Year

Abstract

In this work, we introduce the first benchmark for evaluating the capabilities of large language models (LLMs) in understanding and generating responses in Tunisian Arabic. To achieve this, we construct a dataset of Tunisian Arabic instructions and prompt ten widely-used LLMs that claim to support Arabic. We then assess the LLM responses through both human and LLMbased evaluations across four criteria: quality, correctness, relevance, and dialectal adherence. We analyze the agreement and correlation between these judgments and identify GPT-4o as our automated judge model based on its high correlation with human ratings, and generate a final leaderboard using this model. Our error analysis reveals that most of the LLMs that were evaluated struggle with recognizing and properly responding in Tunisian Arabic. To facilitate further research, we release our dataset, along with gold-standard humanwritten responses for all 744 instructions, and our evaluation framework, allowing others to benchmark their own models.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 5924bceb-c16d-4f05-9bd3-c1cc69b39d64

Builds on6

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines