Enhancing Knowledge Bases with Quantity Facts
Vinh Thinh Ho, Daria Stepanova, Dragan Milchevski, Jannik Strötgen, Gerhard Weikum
Abstract
Machine knowledge about the world’s entities should include quantity properties, such as heights of buildings, running times of athletes, energy efficiency of car models, energy production of power plants, and more. State-of-the-art knowledge bases (KBs), such as Wikidata, cover many relevant entities but often miss the corresponding quantities. Prior work on extracting quantity facts from web contents focused on high precision for top-ranked outputs, but did not tackle the KB coverage issue. This paper presents a recall-oriented approach which aims to close this gap in knowledge-base coverage. Our method is based on iterative learning for extracting quantity facts, with two novel contributions to boost recall for KB augmentation without sacrificing the quality standards of the knowledge base. The first contribution is a query expansion technique to capture a larger pool of fact candidates. The second contribution is a novel technique for harnessing observations on value distributions for self-consistency. Experiments with extractions from more than 13 million web documents demonstrate the benefits of our method.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers3
- Representation Learning on Hyper-Relational and Numeric Knowledge Graphs with TransformersChanyoung Chung, Jaejun Lee, Joyce Jiyoung WhangKDD 2023 · 14 citations
- Learning from Both Structural and Textual Knowledge for Inductive Knowledge Graph CompletionKunxun Qi, Jianfeng Du, Hai WanNeurIPS 2023 · 7 citations
- NumColBERT: Non-Intrusive Numeracy Injection for Late-Interaction Retrieval ModelsHaruki Fujimaki, Makoto P. KatoSIGIR 2026
Related papers
- Extracting Contextualized Quantity Facts from Web TablesVinh Thinh Ho, Koninika Pal, Simon Razniewski, Klaus Berberich et al.WWW 2021 · 13 citations
- Wikidata as a seed for Web ExtractionKunpeng Guo, Dennis Diefenbach, Antoine Gourru, Christophe GravierWWW 2023 · 6 citations
- Open Knowledge Enrichment for Long-tail EntitiesErmei Cao, Difeng Wang, Jiacheng Huang, Wei HuWWW 2020 · 51 citations
- Increasing Coverage and Precision of Textual Information in Multilingual Knowledge GraphsSimone Conia, Min Li, Daniel Lee, Umar Farooq Minhas et al.EMNLP 2023 · 3 citations
- Wiki2Prop: A Multimodal Approach for Predicting Wikidata Properties from WikipediaMichael Luggen, Julien Audiffren, Djellel Eddine Difallah, Philippe Cudré-MaurouxWWW 2021 · 16 citations
