HGNet: Scalable Foundation Model for Automated Knowledge Graph Generation from Scientific Literature
Devvrat Joshi, Islem Rekik
摘要
Automated knowledge graph (KG) construction is essential for navigating the rapidly expanding body of scientific literature. However, existing approaches face persistent challenges: they struggle to recognize long multi-word entities, often fail to generalize across domains, and typically overlook the hierarchical and logically constrained nature of scientific knowledge. While general-purpose large language models (LLMs) offer some adaptability, they are computationally expensive and yield inconsistent accuracy on specialized, domain-heavy tasks such as scientific knowledge graph construction. As a result, current KGs are shallow and inconsistent, limiting their utility for exploration and synthesis. We propose a two-stage framework for scalable, zero-shot scientific KG construction. The first stage, Z-NERD, introduces (i) Orthogonal Semantic Decomposition (OSD), which promotes domain-agnostic entity recognition by isolating semantic "turns" in text, and (ii) a Multi-Scale TCQK attention mechanism that captures coherent multi-word entities through n-gram-aware attention heads. The second stage, HGNet, performs relation extraction with hierarchy-aware message passing, explicitly modeling parent, child, and peer relations. To enforce global consistency, we introduce two complementary objectives: a Differentiable Hierarchy Loss to discourage cycles and shortcut edges, and a Continuum Abstraction Field (CAF) Loss that embeds abstraction levels along a learnable axis in Euclidean space. To the best of our knowledge, this is the first approach to formalize hierarchical abstraction as a continuous property within standard Euclidean embeddings, offering a simpler and more interpretable alternative to hyperbolic methods. To address data scarcity, we also release SPHERE 1 , a large-scale, multidomain benchmark for hierarchical relation extraction. Our framework establishes a new state of the art on benchmarks such as SciERC, SciER and SPHERE benchmarks, improving named entity recognition (NER) by 8.08% and relation extraction (RE) by 5.99% on the official out-of-distribution test sets. In zeroshot settings, the gains are even more pronounced, with improvements of 10.76% for NER and 26.2% for RE, marking a significant step toward reliable and scalable scientific knowledge graph construction. Our HGNet code is available at https://github.com/basiralab/HGNet .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- LUKE: Deep Contextualized Entity Representations with Entity-aware Self-attentionIkuya Yamada, Akari Asai, Hiroyuki Shindo, Hideaki Takeda 等EMNLP 2020 · 被引用 562 次
- Packed Levitated Marker for Entity and Relation ExtractionDeming Ye, Yankai Lin, Peng Li, Maosong SunACL 2022 · 被引用 140 次
- UniversalNER: Targeted Distillation from Large Language Models for Open Named Entity RecognitionWenxuan Zhou, Sheng Zhang, Yu Gu, Muhao Chen 等ICLR 2024 · 被引用 118 次
- Modeling Heterogeneous Hierarchies with Relation-specific Hyperbolic ConesYushi Bai, Zhitao Ying, Hongyu Ren, Jure LeskovecNeurIPS 2021 · 被引用 84 次
- Entity-centered Cross-document Relation ExtractionFengqi Wang, Fei Li, Hao Fei, Jingye Li 等EMNLP 2022 · 被引用 49 次
相关 Paper
- Extract, Define, Canonicalize: An LLM-based Framework for Knowledge Graph ConstructionBowen Zhang, Harold SohEMNLP 2024 · 被引用 65 次
- AgentsKG: A Hierarchical Multi-Agent Framework for Open-Domain Knowledge Graph ConstructionShilong Liu, Yongqiang Liu, Jiye Liu, Xuan Guo 等KDD 2026
- Discovering Latent Facts from Context to Construct Richer Open Knowledge GraphsJinpeng Li, Hang Yu, Ziqi Ma, Peng QiAAAI 2026
- Hyper-KGGen: A Skill-Driven Knowledge Extractor for High-Quality Knowledge Hypergraph GenerationRizhuo Huang, Yifan Feng, Rundong Xue, Shihui Ying 等KDD 2026
- AutoSchemaKG: Autonomous Knowledge Graph Construction through Dynamic Schema Induction from Web-Scale CorporaJiaxin Bai, Wei Fan, Qi Hu, Qing Zong 等ACL 2026 · 被引用 27 次
