Seed-Guided Fine-Grained Entity Typing in Science and Engineering Domains
Yu Zhang, Yunyi Zhang, Yanzhen Shen, Yu Deng, Lucian Popa, Larisa Shwartz, ChengXiang Zhai, Jiawei Han
Abstract
Accurately typing entity mentions from text segments is a fundamental task for various natural language processing applications. Many previous approaches rely on massive human-annotated data to perform entity typing. Nevertheless, collecting such data in highly specialized science and engineering domains (e.g., software engineering and security) can be time-consuming and costly, without mentioning the domain gaps between training and inference data if the model needs to be applied to confidential datasets. In this paper, we study the task of seed-guided fine-grained entity typing in science and engineering domains, which takes the name and a few seed entities for each entity type as the only supervision and aims to classify new entity mentions into both seen and unseen types (i.e., those without seed entities). To solve this problem, we propose SETYPE which first enriches the weak supervision by finding more entities for each seen type from an unlabeled corpus using the contextualized representations of pre-trained language models. It then matches the enriched entities to unlabeled text to get pseudo-labeled samples and trains a textual entailment model that can make inferences for both seen and unseen types. Extensive experiments on two datasets covering four domains demonstrate the effectiveness of SETYPE in comparison with various baselines. Code and data are available at: https://github.com/yuzhimanhua/SEType .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f642db1f-0eb3-4b55-ba2e-174a65cca8e2Builds on14
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Scalable Zero-shot Entity Linking with Dense Entity RetrievalLedell Wu, Fabio Petroni, Martin Josifoski, Sebastian Riedel et al.EMNLP 2020 · 336 citations
- Contextualized Weak Supervision for Text ClassificationDheeraj Mekala, Jingbo ShangACL 2020 · 121 citations
- Empower Entity Set Expansion via Language Model ProbingYunyi Zhang, Jiaming Shen, Jingbo Shang, Jiawei HanACL 2020 · 51 citations
Related papers
- OntoType: Ontology-Guided and Pre-Trained Language Model Assisted Fine-Grained Entity TypingTanay Komarlu, Minhao Jiang, Xuan Wang, Jiawei HanKDD 2024 · 1 citation
- Ontology Enrichment for Effective Fine-grained Entity TypingSiru Ouyang, Jiaxin Huang, Pranav Pillai, Yunyi Zhang et al.KDD 2024 · 5 citations
- Fine-grained Entity Typing without Knowledge BaseJing Qian, Yibin Liu, Lemao Liu, Yangming Li et al.EMNLP 2021 · 1 citation
- Cross-Lingual Contrastive Learning for Fine-Grained Entity Typing for Low-Resource LanguagesXu Han, Yuqi Luo, Weize Chen, Zhiyuan Liu et al.ACL 2022
- Generative Entity Typing with Curriculum LearningSiyu Yuan, Deqing Yang, Jiaqing Liang, Zhixu Li et al.EMNLP 2022 · 11 citations
