Modeling Fine-Grained Entity Types with Box Embeddings
Yasumasa Onoe, Michael Boratko, Andrew McCallum, Greg Durrett
Abstract
Neural entity typing models typically represent fine-grained entity types as vectors in a high-dimensional space, but such spaces are not well-suited to modeling these types' complex interdependencies. We study the ability of box embeddings, which embed concepts as d-dimensional hyperrectangles, to capture hierarchies of types even when these relationships are not defined explicitly in the ontology. Our model represents both types and entity mentions as boxes. Each mention and its context are fed into a BERT-based model to embed that mention in our box space; essentially, this model leverages typological clues present in the surface text to hypothesize a type representation for the mention. Box containment can then be used to derive both the posterior probability of a mention exhibiting a given type and the conditional probability relations between types themselves. We compare our approach with a vector-based typing model and observe state-of-the-art performance on several entity typing benchmarks. In addition to competitive typing performance, our box-based model shows better performance in prediction consistency (predicting a supertype and a subtype together) and confidence (i.e., calibration), demonstrating that the box-based model captures the latent type hierarchies better than the vector-based model does. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers22
- Modeling Label Space Interactions in Multi-label Classification using Box EmbeddingsDhruvesh Patel, Pavitra Dangati, Jay-Yoon Lee, Michael Boratko et al.ICLR 2022 · 28 citations
- A Single Vector Is Not Enough: Taxonomy Expansion via Box EmbeddingsSong Jiang, Qiyue Yao, Qifan Wang, Yizhou SunWWW 2023 · 20 citations
- Transformer-based Entity Typing in Knowledge GraphsZhiwei Hu, Víctor Gutiérrez-Basulto, Zhiliang Xiang, Ru Li et al.EMNLP 2022 · 16 citations
- Divide and Denoise: Learning from Noisy Labels in Fine-Grained Entity Typing with Cluster-Wise Loss CorrectionKunyuan Pang, Haoyu Zhang, Jie Zhou, Ting WangACL 2022 · 13 citations
- Holistic Label Correction for Noisy Multi-Label ClassificationXiaobo Xia, Jiankang Deng, Wei Bao, Yuxuan Du et al.ICCV 2023 · 13 citations
Builds on4
- Query2box: Reasoning over Knowledge Graphs in Vector Space Using Box EmbeddingsHongyu Ren, Weihua Hu, Jure LeskovecICLR 2020 · 355 citations
- BoxE: A Box Embedding Model for Knowledge Base CompletionRalph Abboud, Ismail Ilkan Ceylan, Thomas Lukasiewicz, Tommaso SalvatoriNeurIPS 2020 · 245 citations
- Improving Local Identifiability in Probabilistic Box EmbeddingsShib Sankar Dasgupta, Michael Boratko, Dongxu Zhang, Luke Vilnis et al.NeurIPS 2020 · 75 citations
- Hierarchical Entity Typing via Multi-level Learning to RankTongfei Chen, Yunmo Chen, Benjamin Van DurmeACL 2020 · 51 citations
Related papers
- Improving Entity Linking by Modeling Latent Entity Type InformationShuang Chen, Jinpeng Wang, Feng Jiang, Chin-Yew LinAAAI 2020 · 71 citations
- Polar Ducks and Where to Find Them: Enhancing Entity Linking with Duck Typing and Polar Box EmbeddingsMattia Atzeni, Mikhail Plekhanov, Frédéric A. Dreyer, Nora Kassner et al.EMNLP 2023 · 3 citations
- Binder: Hierarchical Concept Representation through Order Embedding of Binary VectorsCroix Gyurek, Niloy Talukder, Mohammad Al HasanKDD 2024
- A Generate-and-Rank Framework with Semantic Type Regularization for Biomedical Concept NormalizationDongfang Xu, Zeyu Zhang, Steven BethardACL 2020 · 43 citations
- A Supervised Multi-Head Self-Attention Network for Nested Named Entity RecognitionYongxiu Xu, Heyan Huang, Chong Feng, Yue HuAAAI 2021 · 39 citations
