Automated Testing and Improvement of Named Entity Recognition Systems
Boxi Yu, Yiyan Hu, Qiuyang Mang, Wenhan Hu, Pinjia He
Abstract
Named entity recognition (NER) systems have seen rapid progress in recent years due to the development of deep neural networks. These systems are widely used in various natural language processing applications, such as information extraction, question answering, and sentiment analysis. However, the complexity and intractability of deep neural networks can make NER systems unreliable in certain circumstances, resulting in incorrect predictions. For example, NER systems may misidentify female names as chemicals or fail to recognize the names of minority groups, leading to user dissatisfaction. To tackle this problem, we introduce TIN, a novel, widely applicable approach for automatically testing and repairing various NER systems. The key idea for automated testing is that the NER predictions of the same named entities under similar contexts should be identical. The core idea for automated repairing is that similar named entities should have the same NER prediction under the same context. We use TIN to test two SOTA NER models and two commercial NER APIs, i.e., Azure NER and AWS NER. We manually verify 784 of the suspicious issues reported by TIN and find that 702 are erroneous issues, leading to high precision (85.0%-93.4%) across four categories of NER errors: omission, over-labeling, incorrect category, and range error. For automated repairing, TIN achieves a high error reduction rate (26.8%-50.6%) over the four systems under test, which successfully repairs 1,056 out of the 1,877 reported NER errors.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3e4e2757-519b-4fc6-8c04-6b93b93e82a4Cited by top-tier papers2
- Glitch Tokens in Large Language Models: Categorization Taxonomy and Effective DetectionYuxi Li, Yi Liu, Gelei Deng, Ying Zhang et al.FSE 2024 · 12 citations
- Testing Graph Database Systems via Equivalent Query RewritingQiuyang Mang, Aoyang Fang, Boxi Yu, Hanfei Chen et al.ICSE 2024 · 12 citations
Builds on8
- TextBugger: Generating Adversarial Text Against Real-world ApplicationsJinfeng Li, Shouling Ji, Tianyu Du, Bo Li et al.NDSS 2019 · 876 citations
- BERT-ATTACK: Adversarial Attack Against BERT Using BERTLinyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue et al.EMNLP 2020 · 529 citations
- Automatic testing and improvement of machine translationZeyu Sun, Jie M. Zhang, Mark Harman, Mike Papadakis et al.ICSE 2020 · 111 citations
- Structure-invariant testing for machine translationPinjia He, Clara Meister, Zhendong SuICSE 2020 · 84 citations
- Testing Machine Translation via Referential TransparencyPinjia He, Clara Meister, Zhendong SuICSE 2021 · 50 citations
Related papers
- Automated testing of image captioning systemsBoxi Yu, Zhiqing Zhong, Xinran Qin, Jiayi Yao et al.ISSTA 2022 · 24 citations
- Learning to find naming issues with big code and small supervisionJingxuan He, Cheng-Chun Lee, Veselin Raychev, Martin T. VechevPLDI 2021 · 9 citations
- CleanCoNLL: A Nearly Noise-Free Named Entity Recognition DatasetSusanna Rücker, Alan AkbikEMNLP 2023 · 3 citations
- AUTOTRAINER: An Automatic DNN Training Problem Detection and Repair SystemXiaoyu Zhang, Juan Zhai, Shiqing Ma, Chao ShenICSE 2021 · 62 citations
- Editable Neural NetworksAnton Sinitsin, Vsevolod Plokhotnyuk, Dmitry V. Pyrkin, Sergei Popov et al.ICLR 2020 · 210 citations
