Let's Stop Incorrect Comparisons in End-to-end Relation Extraction!
Bruno Taillé, Vincent Guigue, Geoffrey Scoutheeten, Patrick Gallinari
摘要
Despite efforts to distinguish three different evaluation setups (Bekoulis et al., 2018a,b), numerous end-to-end Relation Extraction (RE) articles present unreliable performance comparison to previous work. In this paper, we first identify several patterns of invalid comparisons in published papers and describe them to avoid their propagation. We then propose a small empirical study to quantify the most common mistake's impact and evaluate it leads to overestimating the final RE performance by around 5% on ACE05. We also seize this opportunity to study the unexplored ablations of two recent developments: the use of language model pretraining (specifically BERT) and span-level NER. This meta-analysis emphasizes the need for rigor in the report of both the evaluation setting and the dataset statistics. We finally call for unifying the evaluation setting in end-to-end RE 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Revisiting Relation Extraction in the era of Large Language ModelsSomin Wadhwa, Silvio Amir, Byron C. WallaceACL 2023 · 被引用 145 次
- An Autoregressive Text-to-Graph Framework for Joint Entity and Relation ExtractionUrchade Zaratiana, Nadi Tomeh, Pierre Holat, Thierry CharnoisAAAI 2024 · 被引用 39 次
- REDFM: a Filtered and Multilingual Relation Extraction DatasetPere-Lluís Huguet Cabot, Simone Tedeschi, Axel-Cyrille Ngonga Ngomo, Roberto NavigliACL 2023 · 被引用 9 次
- Unified Structure Generation for Universal Information ExtractionYaojie Lu, Qing Liu, Dai Dai, Xinyan Xiao 等ACL 2022
- UniRE: A Unified Label Space for Entity Relation ExtractionYijun Wang, Changzhi Sun, Yuanbin Wu, Hao Zhou 等ACL 2021
相关 Paper
- On "Scientific Debt" in NLP: A Case for More Rigour in Language Model Pre-Training ResearchMade Nindyatama Nityasya, Haryo Akbarianto Wibowo, Alham Fikri Aji, Genta Indra Winata 等ACL 2023 · 被引用 1 次
- Reproducibility Issues for BERT-based Evaluation MetricsYanran Chen, Jonas Belouadi, Steffen EgerEMNLP 2022 · 被引用 11 次
- OODREB: Benchmarking State-of-the-Art Methods for Out-Of-Distribution Generalization on Relation ExtractionHaotian Chen, Houjing Guo, Bingsheng Chen, Xiangdong ZhouWWW 2024 · 被引用 1 次
- Rogue ScoresMax GruskyACL 2023 · 被引用 11 次
- Do CoNLL-2003 Named Entity Taggers Still Work Well in 2023?Shuheng Liu, Alan RitterACL 2023 · 被引用 9 次
