Interpretable Multi-dataset Evaluation for Named Entity Recognition
Jinlan Fu, Pengfei Liu, Graham Neubig
摘要
With the proliferation of models for natural language processing tasks, it is even harder to understand the differences between models and their relative merits. Simply looking at differences between holistic metrics such as accuracy, BLEU, or F1 does not tell us why or how particular methods perform differently and how diverse datasets influence the model design choices. In this paper, we present a general methodology for interpretable evaluation for the named entity recognition (NER) task. The proposed evaluation method enables us to interpret the differences in models and datasets, as well as the interplay between them, identifying the strengths and weaknesses of current systems. By making our analysis tool available, we make it easy for future researchers to run similar analyses and drive progress in this area: https: //github.com/neulab/InterpretEval.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Polyglot Prompt: Multilingual Multitask Prompt TrainingJinlan Fu, See-Kiong Ng, Pengfei LiuEMNLP 2022 · 被引用 11 次
- XTREME-R: Towards More Challenging and Nuanced Multilingual EvaluationSebastian Ruder, Noah Constant, Jan A. Botha, Aditya Siddhant 等EMNLP 2021 · 被引用 10 次
- SpanNER: Named Entity Re-/Recognition as Span PredictionJinlan Fu, Xuanjing Huang, Pengfei LiuACL 2021
- Temporal Relation Extraction in Clinical Texts: A Span-based Graph Transformer ApproachRochana Chaturvedi, Peyman Baghershahi, Sourav Medya, Barbara Di EugenioACL 2025
它引用的顶会 Paper4
- A Unified MRC Framework for Named Entity RecognitionXiaoya Li, Jingrong Feng, Yuxian Meng, Qinghong Han 等ACL 2020 · 被引用 617 次
- Hierarchical Contextualized Representation for Named Entity RecognitionYing Luo, Fengshun Xiao, Hai ZhaoAAAI 2020 · 被引用 138 次
- Rethinking Generalization of Neural Models: A Named Entity Recognition Case StudyJinlan Fu, Pengfei Liu, Qi ZhangAAAI 2020 · 被引用 79 次
- Leveraging Multi-Token Entities in Document-Level Named Entity RecognitionAnwen Hu, Zhicheng Dou, Jian-Yun Nie, Ji-Rong WenAAAI 2020 · 被引用 24 次
相关 Paper
- DecompEval: Evaluating Generated Texts as Unsupervised Decomposed Question AnsweringPei Ke, Fei Huang, Fei Mi, Yasheng Wang 等ACL 2023 · 被引用 2 次
- Evaluating Neuron Interpretation Methods of NLP ModelsYimin Fan, Fahim Dalvi, Nadir Durrani, Hassan SajjadNeurIPS 2023 · 被引用 11 次
- ERASER: A Benchmark to Evaluate Rationalized NLP ModelsJay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric P. Lehman 等ACL 2020 · 被引用 36 次
- Interpretable NLG for Task-oriented Dialogue Systems with Heterogeneous Rendering MachinesYangming Li, Kaisheng YaoAAAI 2021 · 被引用 4 次
- A Fair and In-Depth Evaluation of Existing End-to-End Entity Linking SystemsHannah Bast, Matthias Hertel, Natalie PrangeEMNLP 2023 · 被引用 3 次
