Globally Aware Contextual Embeddings for Named Entity Recognition in Social Media Streams
Satadisha Saha Bhowmick, Eduard C. Dragut, Weiyi Meng
摘要
An important task for Information Extraction from Microblogs is Named Entity Recognition (NER) that extracts mentions of real-world entities from microblog messages and meta-information like entity type for better entity characterization. A lot of microblog NER systems have rightly sought to prioritize modeling the non-literary nature of microblog text. These systems are trained on offline static datasets and extract a combination of surface-level features – orthographic, lexical, and semantic – from individual messages for noisy text modeling and entity extraction. But given the constantly evolving nature of microblog streams, detecting all entity mentions from such varying yet limited context in short messages remains a difficult problem to generalize. In this paper, we propose the NER Globalizer pipeline better suited for NER on microblog streams. It characterizes the isolated message processing by existing NER systems as modeling local contextual embeddings, where learned knowledge from the immediate context of a message is used to suggest seed entity candidates. Additionally, it recognizes that messages within a microblog stream are topically related and often repeat mentions of the same entity. This suggests building NER systems that go beyond localized processing. By leveraging occurrence mining, the proposed system therefore follows up traditional NER modeling by extracting additional mentions of seed entity candidates that were previously missed. Candidate mentions are separated into well-defined clusters which are then used to generate a pooled global embedding drawn from the collective context of the candidate within a stream. The global embeddings are utilized to separate false positives from entities whose mentions are produced in the final NER output. Our experiments show that the proposed NER system exhibits superior effectiveness on multiple NER datasets with an average Macro F1 improvement of 47.04% over the best NER baseline while adding only a small computational overhead.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper5
- SciER: An Entity and Relation Extraction Dataset for Datasets, Methods, and Tasks in Scientific DocumentsQi Zhang, Zhijia Chen, Huitong Pan, Cornelia Caragea 等EMNLP 2024 · 被引用 7 次
- Mitigating Data Sparsity in Integrated Data through Text ConceptualizationMd. Ataur Rahman, Sergi Nadal, Oscar Romero, Dimitris SacharidisICDE 2024 · 被引用 3 次
- LLMLog: Advanced Log Template Generation via LLM-driven Multi-Round AnnotationFei Teng, Haoyang Li, Lei ChenVLDB 2025 · 被引用 2 次
- ReliK: A Reliability Measure for Knowledge Graph EmbeddingsMaximilian K. Egger, Wenyue Ma, Davide Mottin, Panagiotis Karras 等WWW 2024 · 被引用 2 次
- HYPPO: Using Equivalences to Optimize Pipelines in Exploratory Machine LearningAntonios Kontaxakis, Dimitris Sacharidis, Alkis Simitsis, Alberto Abelló 等ICDE 2024 · 被引用 1 次
相关 Paper
- Boosting Entity Mention Detection for Targetted Twitter Streams with Global Contextual EmbeddingsSatadisha Saha Bhowmick, Eduard C. Dragut, Weiyi MengICDE 2022 · 被引用 3 次
- Hierarchical Aligned Multimodal Learning for NER on Tweet PostsPeipei Liu, Hong Li, Yimo Ren, Jie Liu 等AAAI 2024 · 被引用 11 次
- Fine-Grained Multimodal Named Entity Recognition and Grounding with a Generative FrameworkJieming Wang, Ziyan Li, Jianfei Yu, Li Yang 等ACM MM 2023 · 被引用 11 次
- Grounded Multimodal Named Entity Recognition on Social MediaJianfei Yu, Ziyan Li, Jieming Wang, Rui XiaACL 2023 · 被引用 32 次
- Leveraging Multi-Token Entities in Document-Level Named Entity RecognitionAnwen Hu, Zhicheng Dou, Jian-Yun Nie, Ji-Rong WenAAAI 2020 · 被引用 24 次
