Let's Stop Incorrect Comparisons in End-to-end Relation Extraction!
Bruno Taillé, Vincent Guigue, Geoffrey Scoutheeten, Patrick Gallinari
Abstract
Despite efforts to distinguish three different evaluation setups (Bekoulis et al., 2018a,b), numerous end-to-end Relation Extraction (RE) articles present unreliable performance comparison to previous work. In this paper, we first identify several patterns of invalid comparisons in published papers and describe them to avoid their propagation. We then propose a small empirical study to quantify the most common mistake's impact and evaluate it leads to overestimating the final RE performance by around 5% on ACE05. We also seize this opportunity to study the unexplored ablations of two recent developments: the use of language model pretraining (specifically BERT) and span-level NER. This meta-analysis emphasizes the need for rigor in the report of both the evaluation setting and the dataset statistics. We finally call for unifying the evaluation setting in end-to-end RE 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 41fe9468-ce9b-498f-8475-0ce887258dc1Cited by top-tier papers5
- Revisiting Relation Extraction in the era of Large Language ModelsSomin Wadhwa, Silvio Amir, Byron C. WallaceACL 2023 · 145 citations
- An Autoregressive Text-to-Graph Framework for Joint Entity and Relation ExtractionUrchade Zaratiana, Nadi Tomeh, Pierre Holat, Thierry CharnoisAAAI 2024 · 39 citations
- REDFM: a Filtered and Multilingual Relation Extraction DatasetPere-Lluís Huguet Cabot, Simone Tedeschi, Axel-Cyrille Ngonga Ngomo, Roberto NavigliACL 2023 · 9 citations
- Unified Structure Generation for Universal Information ExtractionYaojie Lu, Qing Liu, Dai Dai, Xinyan Xiao et al.ACL 2022
- UniRE: A Unified Label Space for Entity Relation ExtractionYijun Wang, Changzhi Sun, Yuanbin Wu, Hao Zhou et al.ACL 2021
Related papers
- On "Scientific Debt" in NLP: A Case for More Rigour in Language Model Pre-Training ResearchMade Nindyatama Nityasya, Haryo Akbarianto Wibowo, Alham Fikri Aji, Genta Indra Winata et al.ACL 2023 · 1 citation
- Reproducibility Issues for BERT-based Evaluation MetricsYanran Chen, Jonas Belouadi, Steffen EgerEMNLP 2022 · 11 citations
- OODREB: Benchmarking State-of-the-Art Methods for Out-Of-Distribution Generalization on Relation ExtractionHaotian Chen, Houjing Guo, Bingsheng Chen, Xiangdong ZhouWWW 2024 · 1 citation
- Rogue ScoresMax GruskyACL 2023 · 11 citations
- Do CoNLL-2003 Named Entity Taggers Still Work Well in 2023?Shuheng Liu, Alan RitterACL 2023 · 9 citations
