Document-to-Database: Extraction Meets Relational Semantics
Zhengxuan Zhang, Zhuowen Liang, Jiazhuo Chen, Haixun Wang, Nan Tang
Abstract
Bridging the gap between unstructured documents and relational databases is challenging because document extraction operates locally, whereas databases enforce global semantics through schemas, keys, and constraints. Consequently, existing one-shot large language model (LLM) extraction approaches often fail to reconcile results with relational semantics, yielding inconsistent and hard-to-audit outputs. We present DataMosaic, a document-to-database (Doc2DB) system that explicitly mediates between extraction and database semantics. Given a database schema and constraints, a central orchestrator coordinates entity and relationship extraction alongside verification, repair, and targeted re-extraction within a closed extract-verify-iterate loop. By systematically resolving document ambiguity and constraint violations, DataMosaic incrementally constructs accurate and semantically consistent databases. Featuring pluggable extractors, verifiers and repair operators, experiments in diverse datasets show that DataMosaic substantially reduces constraint violations and improves database-level accuracy over strong Doc2DB baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7fdddc55-d6f2-4238-a74c-6ae20bbcf27cBuilds on11
- LayoutLM: Pre-training of Text and Layout for Document Image UnderstandingYiheng Xu, Minghao Li, Lei Cui, Shaohan Huang et al.KDD 2020 · 575 citations
- DocETL: Agentic Query Rewriting and Evaluation for Complex Document ProcessingShreya Shankar, Tristan Chambers, Tarak Shah, Aditya G. Parameswaran et al.VLDB 2025 · 62 citations
- HAIPipe: Combining Human-generated and Machine-generated Pipelines for Data PreparationSibei Chen, Nan Tang, Ju Fan, Xuemi Yan et al.SIGMOD 2023 · 25 citations
- VisJudge-Bench: Aesthetics and Quality Assessment of VisualizationsYupeng Xie, Zhiyang Zhang, Yifan Wu, Sirong Lu et al.ICLR 2026 · 23 citations
- AutoPrep: Natural Language Question-Aware Data Preparation with a Multi-Agent FrameworkMeihao Fan, Ju Fan, Nan Tang, Lei Cao et al.VLDB 2025 · 10 citations
Related papers
- SQUiD: Synthesizing Relational Databases from Unstructured TextMushtari Sadia, Zhenning Yang, Yunming Xiao, Ang Chen et al.EMNLP 2025
- LinkAlign: Scalable Schema Linking for Real-World Large-Scale Multi-Database Text-to-SQLYihan Wang, Peiyu Liu, Xin YangEMNLP 2025 · 2 citations
- DBugScribe: Automatic Database Bug Reproduction from Community ReportsSuyang Zhong, Mo Sha, Sheng Wang, Fangyuan Zhou et al.SIGMOD 2026 · 1 citation
- SchemaRAG: A Schema-aware Retrieval-Augmented Generation Framework for Text-to-SQLDi Wu, Zetong Tang, Yi He, Xin LuoSIGMOD 2026 · 9 citations
- Doctopus: Budget-aware Structural Table Extraction from Unstructured DocumentsChengliang Chai, Jiajun Li, Yuhao Deng, Yuanhao Zhong et al.VLDB 2025 · 5 citations
