REDFM: a Filtered and Multilingual Relation Extraction Dataset
Pere-Lluís Huguet Cabot, Simone Tedeschi, Axel-Cyrille Ngonga Ngomo, Roberto Navigli
Abstract
Relation Extraction (RE) is a task that identifies relationships between entities in a text, enabling the acquisition of relational facts and bridging the gap between natural language and structured knowledge. However, current RE models often rely on small datasets with low coverage of relation types, particularly when working with languages other than English. In this paper, we address the above issue and provide two new resources that enable the training and evaluation of multilingual RE systems. First, we present SRED FM , an automatically annotated dataset covering 18 languages, 400 relation types, 13 entity types, totaling more than 40 million triplet instances. Second, we propose RED FM , a smaller, human-revised dataset for seven languages that allows for the evaluation of multilingual RE systems. To demonstrate the utility of these novel datasets, we experiment with the first end-to-end multilingual RE model, mREBEL, that extracts triplets, including entity types, in multiple languages. We release our resources and model checkpoints at https://www.github.com/babelscape/rebel .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e91dc878-7fe9-4329-8fca-2c04c615738eCited by top-tier papers3
- WojoodRelations: Arabic Relation Extraction Corpus and ModelingAlaa Aljabari, Mohammed Khalilia, Mustafa JarrarEMNLP 2025 · 1 citation
- Translation and Fusion Improves Cross-lingual Information ExtractionYang Chen, Vedaant Shah, Alan RitterACL 2025
- Towards Fast and Accurate Modeling for Cross-Lingual Label ProjectionThang Le, Huy Huu Nguyen, Anh Tuan Luu, Thamar Solorio et al.ACL 2026
Builds on6
- LUKE: Deep Contextualized Entity Representations with Entity-aware Self-attentionIkuya Yamada, Akari Asai, Hiroyuki Shindo, Hideaki Takeda et al.EMNLP 2020 · 562 citations
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 394 citations
- Two are Better than One: Joint Entity and Relation Extraction with Table-Sequence EncodersJue Wang, Wei LuEMNLP 2020 · 209 citations
- Re-TACRED: Addressing Shortcomings of the TACRED DatasetGeorge Stoica, Emmanouil Antonios Platanios, Barnabás PóczosAAAI 2021 · 146 citations
- Multilingual Relation Classification via Efficient and Effective PromptingYuxuan Chen, David Harbecke, Leonhard HennigEMNLP 2022 · 13 citations
Related papers
- MultiTACRED: A Multilingual Version of the TAC Relation Extraction DatasetLeonhard Hennig, Philippe Thomas, Sebastian MöllerACL 2023 · 4 citations
- CodRED: A Cross-Document Relation Extraction Dataset for Acquiring Knowledge in the WildYuan Yao, Jiaju Du, Yankai Lin, Peng Li et al.EMNLP 2021 · 18 citations
- HistRED: A Historical Document-Level Relation Extraction DatasetSoyoung Yang, Minseok Choi, Youngwoo Cho, Jaegul ChooACL 2023 · 9 citations
- MEE: A Novel Multilingual Event Extraction DatasetAmir Pouran Ben Veyseh, Javid Ebrahimi, Franck Dernoncourt, Thien Huu NguyenEMNLP 2022 · 3 citations
- Entity Linking in 100 LanguagesJan A. Botha, Zifei Shan, Daniel GillickEMNLP 2020
