Why Attention? Analyze BiLSTM Deficiency and Its Remedies in the Case of NER
Peng-Hsuan Li, Tsu-Jui Fu, Wei-Yun Ma
Abstract
BiLSTM has been prevalently used as a core module for NER in a sequence-labeling setup. State-of-the-art approaches use BiLSTM with additional resources such as gazetteers, language-modeling, or multi-task supervision to further improve NER. This paper instead takes a step back and focuses on analyzing problems of BiLSTM itself and how exactly self-attention can bring improvements. We formally show the limitation of (CRF-)BiLSTM in modeling cross-context patterns for each word -the XOR limitation. Then, we show that two types of simple cross-structures -self-attention and Cross-BiLSTM -can effectively remedy the problem. We test the practical impacts of the deficiency on real-world NER datasets, OntoNotes 5.0 and WNUT 2017, with clear and consistent improvements over the baseline, up to 8.7% on some of the multi-token entity mentions. We give in-depth analyses of the improvements across several aspects of NER, especially the identification of multi-token mentions. This study should lay a sound foundation for future improvements on sequence-labeling NER 1 . Introduction Named Entity Recognition (NER) is a core task for information extraction. Originally a structured prediction task, NER has since been formulated as a task of sequential token labeling. BiLSTM-CNN uses a CNN to encode each word and then uses bi-directional LSTMs to encode past and future context respectively at each time step. With stateof-the-art empirical results, most regard it as a robust core module for sequence-labeling NER (
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Related papers
- Hierarchical Contextualized Representation for Named Entity RecognitionYing Luo, Fengshun Xiao, Hai ZhaoAAAI 2020 · 138 citations
- Leveraging Multi-Token Entities in Document-Level Named Entity RecognitionAnwen Hu, Zhicheng Dou, Jian-Yun Nie, Ji-Rong WenAAAI 2020 · 24 citations
- Attention and Edge-Label Guided Graph Convolutional Networks for Named Entity RecognitionRenjie Zhou, Zhongyi Xie, Jian Wan, Jilin Zhang et al.EMNLP 2022 · 5 citations
- A Supervised Multi-Head Self-Attention Network for Nested Named Entity RecognitionYongxiu Xu, Heyan Huang, Chong Feng, Yue HuAAAI 2021 · 39 citations
- HIT: Nested Named Entity Recognition via Head-Tail Pair and Token InteractionYu Wang, Yun Li, Hanghang Tong, Ziye ZhuEMNLP 2020 · 36 citations
