Interpretable Multi-dataset Evaluation for Named Entity Recognition
Jinlan Fu, Pengfei Liu, Graham Neubig
Abstract
With the proliferation of models for natural language processing tasks, it is even harder to understand the differences between models and their relative merits. Simply looking at differences between holistic metrics such as accuracy, BLEU, or F1 does not tell us why or how particular methods perform differently and how diverse datasets influence the model design choices. In this paper, we present a general methodology for interpretable evaluation for the named entity recognition (NER) task. The proposed evaluation method enables us to interpret the differences in models and datasets, as well as the interplay between them, identifying the strengths and weaknesses of current systems. By making our analysis tool available, we make it easy for future researchers to run similar analyses and drive progress in this area: https: //github.com/neulab/InterpretEval.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cf95faf9-cb53-42c5-a7d1-594f811364bbCited by top-tier papers4
- Polyglot Prompt: Multilingual Multitask Prompt TrainingJinlan Fu, See-Kiong Ng, Pengfei LiuEMNLP 2022 · 11 citations
- XTREME-R: Towards More Challenging and Nuanced Multilingual EvaluationSebastian Ruder, Noah Constant, Jan A. Botha, Aditya Siddhant et al.EMNLP 2021 · 10 citations
- SpanNER: Named Entity Re-/Recognition as Span PredictionJinlan Fu, Xuanjing Huang, Pengfei LiuACL 2021
- Temporal Relation Extraction in Clinical Texts: A Span-based Graph Transformer ApproachRochana Chaturvedi, Peyman Baghershahi, Sourav Medya, Barbara Di EugenioACL 2025
Builds on4
- A Unified MRC Framework for Named Entity RecognitionXiaoya Li, Jingrong Feng, Yuxian Meng, Qinghong Han et al.ACL 2020 · 617 citations
- Hierarchical Contextualized Representation for Named Entity RecognitionYing Luo, Fengshun Xiao, Hai ZhaoAAAI 2020 · 138 citations
- Rethinking Generalization of Neural Models: A Named Entity Recognition Case StudyJinlan Fu, Pengfei Liu, Qi ZhangAAAI 2020 · 79 citations
- Leveraging Multi-Token Entities in Document-Level Named Entity RecognitionAnwen Hu, Zhicheng Dou, Jian-Yun Nie, Ji-Rong WenAAAI 2020 · 24 citations
Related papers
- DecompEval: Evaluating Generated Texts as Unsupervised Decomposed Question AnsweringPei Ke, Fei Huang, Fei Mi, Yasheng Wang et al.ACL 2023 · 2 citations
- Evaluating Neuron Interpretation Methods of NLP ModelsYimin Fan, Fahim Dalvi, Nadir Durrani, Hassan SajjadNeurIPS 2023 · 11 citations
- ERASER: A Benchmark to Evaluate Rationalized NLP ModelsJay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric P. Lehman et al.ACL 2020 · 36 citations
- Interpretable NLG for Task-oriented Dialogue Systems with Heterogeneous Rendering MachinesYangming Li, Kaisheng YaoAAAI 2021 · 4 citations
- A Fair and In-Depth Evaluation of Existing End-to-End Entity Linking SystemsHannah Bast, Matthias Hertel, Natalie PrangeEMNLP 2023 · 3 citations
