Globally Aware Contextual Embeddings for Named Entity Recognition in Social Media Streams
Satadisha Saha Bhowmick, Eduard C. Dragut, Weiyi Meng
Abstract
An important task for Information Extraction from Microblogs is Named Entity Recognition (NER) that extracts mentions of real-world entities from microblog messages and meta-information like entity type for better entity characterization. A lot of microblog NER systems have rightly sought to prioritize modeling the non-literary nature of microblog text. These systems are trained on offline static datasets and extract a combination of surface-level features – orthographic, lexical, and semantic – from individual messages for noisy text modeling and entity extraction. But given the constantly evolving nature of microblog streams, detecting all entity mentions from such varying yet limited context in short messages remains a difficult problem to generalize. In this paper, we propose the NER Globalizer pipeline better suited for NER on microblog streams. It characterizes the isolated message processing by existing NER systems as modeling local contextual embeddings, where learned knowledge from the immediate context of a message is used to suggest seed entity candidates. Additionally, it recognizes that messages within a microblog stream are topically related and often repeat mentions of the same entity. This suggests building NER systems that go beyond localized processing. By leveraging occurrence mining, the proposed system therefore follows up traditional NER modeling by extracting additional mentions of seed entity candidates that were previously missed. Candidate mentions are separated into well-defined clusters which are then used to generate a pooled global embedding drawn from the collective context of the candidate within a stream. The global embeddings are utilized to separate false positives from entities whose mentions are produced in the final NER output. Our experiments show that the proposed NER system exhibits superior effectiveness on multiple NER datasets with an average Macro F1 improvement of 47.04% over the best NER baseline while adding only a small computational overhead.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers5
- SciER: An Entity and Relation Extraction Dataset for Datasets, Methods, and Tasks in Scientific DocumentsQi Zhang, Zhijia Chen, Huitong Pan, Cornelia Caragea et al.EMNLP 2024 · 7 citations
- Mitigating Data Sparsity in Integrated Data through Text ConceptualizationMd. Ataur Rahman, Sergi Nadal, Oscar Romero, Dimitris SacharidisICDE 2024 · 3 citations
- LLMLog: Advanced Log Template Generation via LLM-driven Multi-Round AnnotationFei Teng, Haoyang Li, Lei ChenVLDB 2025 · 2 citations
- ReliK: A Reliability Measure for Knowledge Graph EmbeddingsMaximilian K. Egger, Wenyue Ma, Davide Mottin, Panagiotis Karras et al.WWW 2024 · 2 citations
- HYPPO: Using Equivalences to Optimize Pipelines in Exploratory Machine LearningAntonios Kontaxakis, Dimitris Sacharidis, Alkis Simitsis, Alberto Abelló et al.ICDE 2024 · 1 citation
Related papers
- Boosting Entity Mention Detection for Targetted Twitter Streams with Global Contextual EmbeddingsSatadisha Saha Bhowmick, Eduard C. Dragut, Weiyi MengICDE 2022 · 3 citations
- Hierarchical Aligned Multimodal Learning for NER on Tweet PostsPeipei Liu, Hong Li, Yimo Ren, Jie Liu et al.AAAI 2024 · 11 citations
- Fine-Grained Multimodal Named Entity Recognition and Grounding with a Generative FrameworkJieming Wang, Ziyan Li, Jianfei Yu, Li Yang et al.ACM MM 2023 · 11 citations
- Grounded Multimodal Named Entity Recognition on Social MediaJianfei Yu, Ziyan Li, Jieming Wang, Rui XiaACL 2023 · 32 citations
- Leveraging Multi-Token Entities in Document-Level Named Entity RecognitionAnwen Hu, Zhicheng Dou, Jian-Yun Nie, Ji-Rong WenAAAI 2020 · 24 citations
