Lune

EMNLP2025顶会

TounsiBench: Benchmarking Large Language Models for Tunisian Arabic

Souha Hassine, Asma Arrak, Marouene Addhoum, Steven R. Wilson

2025年份

摘要

In this work, we introduce the first benchmark for evaluating the capabilities of large language models (LLMs) in understanding and generating responses in Tunisian Arabic. To achieve this, we construct a dataset of Tunisian Arabic instructions and prompt ten widely-used LLMs that claim to support Arabic. We then assess the LLM responses through both human and LLMbased evaluations across four criteria: quality, correctness, relevance, and dialectal adherence. We analyze the agreement and correlation between these judgments and identify GPT-4o as our automated judge model based on its high correlation with human ratings, and generate a final leaderboard using this model. Our error analysis reveals that most of the LLMs that were evaluated struggle with recognizing and properly responding in Tunisian Arabic. To facilitate further research, we release our dataset, along with gold-standard humanwritten responses for all 744 instructions, and our evaluation framework, allowing others to benchmark their own models.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 5924bceb-c16d-4f05-9bd3-c1cc69b39d64

它引用的顶会 Paper6

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖