REDFM: a Filtered and Multilingual Relation Extraction Dataset
Pere-Lluís Huguet Cabot, Simone Tedeschi, Axel-Cyrille Ngonga Ngomo, Roberto Navigli
摘要
Relation Extraction (RE) is a task that identifies relationships between entities in a text, enabling the acquisition of relational facts and bridging the gap between natural language and structured knowledge. However, current RE models often rely on small datasets with low coverage of relation types, particularly when working with languages other than English. In this paper, we address the above issue and provide two new resources that enable the training and evaluation of multilingual RE systems. First, we present SRED FM , an automatically annotated dataset covering 18 languages, 400 relation types, 13 entity types, totaling more than 40 million triplet instances. Second, we propose RED FM , a smaller, human-revised dataset for seven languages that allows for the evaluation of multilingual RE systems. To demonstrate the utility of these novel datasets, we experiment with the first end-to-end multilingual RE model, mREBEL, that extracts triplets, including entity types, in multiple languages. We release our resources and model checkpoints at https://www.github.com/babelscape/rebel .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- WojoodRelations: Arabic Relation Extraction Corpus and ModelingAlaa Aljabari, Mohammed Khalilia, Mustafa JarrarEMNLP 2025 · 被引用 1 次
- Translation and Fusion Improves Cross-lingual Information ExtractionYang Chen, Vedaant Shah, Alan RitterACL 2025
- Towards Fast and Accurate Modeling for Cross-Lingual Label ProjectionThang Le, Huy Huu Nguyen, Anh Tuan Luu, Thamar Solorio 等ACL 2026
它引用的顶会 Paper6
- LUKE: Deep Contextualized Entity Representations with Entity-aware Self-attentionIkuya Yamada, Akari Asai, Hiroyuki Shindo, Hideaki Takeda 等EMNLP 2020 · 被引用 562 次
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 被引用 394 次
- Two are Better than One: Joint Entity and Relation Extraction with Table-Sequence EncodersJue Wang, Wei LuEMNLP 2020 · 被引用 209 次
- Re-TACRED: Addressing Shortcomings of the TACRED DatasetGeorge Stoica, Emmanouil Antonios Platanios, Barnabás PóczosAAAI 2021 · 被引用 146 次
- Multilingual Relation Classification via Efficient and Effective PromptingYuxuan Chen, David Harbecke, Leonhard HennigEMNLP 2022 · 被引用 13 次
相关 Paper
- MultiTACRED: A Multilingual Version of the TAC Relation Extraction DatasetLeonhard Hennig, Philippe Thomas, Sebastian MöllerACL 2023 · 被引用 4 次
- CodRED: A Cross-Document Relation Extraction Dataset for Acquiring Knowledge in the WildYuan Yao, Jiaju Du, Yankai Lin, Peng Li 等EMNLP 2021 · 被引用 18 次
- HistRED: A Historical Document-Level Relation Extraction DatasetSoyoung Yang, Minseok Choi, Youngwoo Cho, Jaegul ChooACL 2023 · 被引用 9 次
- MEE: A Novel Multilingual Event Extraction DatasetAmir Pouran Ben Veyseh, Javid Ebrahimi, Franck Dernoncourt, Thien Huu NguyenEMNLP 2022 · 被引用 3 次
- Entity Linking in 100 LanguagesJan A. Botha, Zifei Shan, Daniel GillickEMNLP 2020
