Lune

EMNLP2025Top-tier venue

WojoodRelations: Arabic Relation Extraction Corpus and Modeling

Alaa Aljabari, Mohammed Khalilia, Mustafa Jarrar

2025Year
1Citations

Abstract

Relation extraction (RE) is a core task in natural language processing, crucial for semantic understanding, knowledge graph construction, and enhancing downstream applications. Existing work on Arabic RE remains limited due to the language's rich morphology and syntactic complexity, and the lack of large, highquality datasets. In this paper, we present Wojood Relations , the largest and most diverse Arabic RE corpus to date, containing over 33K sentences (∼ 550K tokens) annotated with ∼ 15K relation triples across 40 relation types. The corpus is built on top of Wojood NER dataset with manual relation annotations carried out by expert annotators, achieving a Cohen's κ of 0.92, indicating high reliability. In addition, we propose two methods: NLI-RE, which formulates RE as a binary natural language inference problem using relation-aware templates, and GPT-Joint, a few-shot LLM framework for joint entity and RE via relationaware retrieval. Finally, we benchmark the dataset using both supervised models and incontext learning with LLMs. Supervised models achieve 92.89% F1 for RE, while LLMs obtain 72.73% F1 for joint entity and RE. These results establish strong baselines, highlight key challenges, and provide a foundation for advancing Arabic RE research.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 9f706cfa-378e-4eed-8453-040e037e1c3a

Builds on7

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines