OpenForge: Probabilistic Metadata Integration
Tianji Cong, Fatemeh Nargesian, Junjie Xing, H. V. Jagadish
Abstract
Modern data stores increasingly rely on metadata for enabling diverse activities such as data cataloging and search. However, metadata curation remains a labor-intensive task, and the broader challenge of metadata maintenance-ensuring its consistency, usefulness, and freshness-has been largely overlooked. In this work, we tackle the problem of resolving relationships among metadata concepts from disparate sources. These relationships are critical for creating clean, consistent, and up-to-date metadata repositories, and a central challenge for metadata integration. We propose OpenForge, a two-stage prior-posterior framework for metadata integration. In the first stage, OpenForge exploits multiple methods including fine-tuned large language models to obtain prior beliefs about concept relationships. In the second stage, OpenForge refines these predictions by leveraging Markov Random Field, a probabilistic graphical model. We formalize metadata integration as an optimization problem, where the objective is to identify the relationship assignments that maximize the joint probability of assignments. The MRF formulation allows OpenForge to capture prior beliefs while encoding critical relationship properties, such as transitivity, in probabilistic inference. Experiments on real-world datasets demonstrate the effectiveness and efficiency of OpenForge. On a use case of matching two metadata vocabularies, OpenForge outperforms GPT-4, the second-best method, by 25 F1-score points.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ad861053-4cb7-4bb4-a557-e16116a21657Builds on9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Dataset Discovery in Data LakesAlex Bogatu, Alvaro A. A. Fernandes, Norman W. Paton, Nikolaos KonstantinouICDE 2020 · 118 citations
- Annotating Columns with Pre-trained Language ModelsYoshihiko Suhara, Jinfeng Li, Yuliang Li, Dan Zhang et al.SIGMOD 2022 · 81 citations
- SANTOS: Relationship-based Semantic Table Union SearchAamod Khatiwada, Grace Fan, Roee Shraga, Zixuan Chen et al.SIGMOD 2023 · 61 citations
Related papers
- Paths-over-Graph: Knowledge Graph Empowered Large Language Model ReasoningXingyu Tan, Xiaoyang Wang, Qing Liu, Xiwei Xu et al.WWW 2025 · 86 citations
- Topic-Oriented Open Relation Extraction with A Priori Seed GenerationLinyi Ding, Jinfeng Xiao, Sizhe Zhou, Chaoqi Yang et al.EMNLP 2024 · 2 citations
- MRF-Chat: Improving Dialogue with Markov Random FieldsIshaan Grover, Matthew Huggins, Cynthia Breazeal, Hae Won ParkEMNLP 2021
- From Assistant to Independent Developer — Are GPTs Ready for Software Development?Dezhi Ran, Yuan Cao, Mengzhou Wu, Simin Chen et al.ICLR 2026 · 4 citations
- Enhancing Self-Consistency and Performance of Pre-Trained Language Models through Natural Language InferenceEric Mitchell, Joseph J. Noh, Siyan Li, William S. Armstrong et al.EMNLP 2022 · 18 citations
