Document-to-Database: Extraction Meets Relational Semantics
Zhengxuan Zhang, Zhuowen Liang, Jiazhuo Chen, Haixun Wang, Nan Tang
摘要
Bridging the gap between unstructured documents and relational databases is challenging because document extraction operates locally, whereas databases enforce global semantics through schemas, keys, and constraints. Consequently, existing one-shot large language model (LLM) extraction approaches often fail to reconcile results with relational semantics, yielding inconsistent and hard-to-audit outputs. We present DataMosaic, a document-to-database (Doc2DB) system that explicitly mediates between extraction and database semantics. Given a database schema and constraints, a central orchestrator coordinates entity and relationship extraction alongside verification, repair, and targeted re-extraction within a closed extract-verify-iterate loop. By systematically resolving document ambiguity and constraint violations, DataMosaic incrementally constructs accurate and semantically consistent databases. Featuring pluggable extractors, verifiers and repair operators, experiments in diverse datasets show that DataMosaic substantially reduces constraint violations and improves database-level accuracy over strong Doc2DB baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- LayoutLM: Pre-training of Text and Layout for Document Image UnderstandingYiheng Xu, Minghao Li, Lei Cui, Shaohan Huang 等KDD 2020 · 被引用 575 次
- DocETL: Agentic Query Rewriting and Evaluation for Complex Document ProcessingShreya Shankar, Tristan Chambers, Tarak Shah, Aditya G. Parameswaran 等VLDB 2025 · 被引用 62 次
- HAIPipe: Combining Human-generated and Machine-generated Pipelines for Data PreparationSibei Chen, Nan Tang, Ju Fan, Xuemi Yan 等SIGMOD 2023 · 被引用 25 次
- VisJudge-Bench: Aesthetics and Quality Assessment of VisualizationsYupeng Xie, Zhiyang Zhang, Yifan Wu, Sirong Lu 等ICLR 2026 · 被引用 23 次
- AutoPrep: Natural Language Question-Aware Data Preparation with a Multi-Agent FrameworkMeihao Fan, Ju Fan, Nan Tang, Lei Cao 等VLDB 2025 · 被引用 10 次
相关 Paper
- SQUiD: Synthesizing Relational Databases from Unstructured TextMushtari Sadia, Zhenning Yang, Yunming Xiao, Ang Chen 等EMNLP 2025
- LinkAlign: Scalable Schema Linking for Real-World Large-Scale Multi-Database Text-to-SQLYihan Wang, Peiyu Liu, Xin YangEMNLP 2025 · 被引用 2 次
- DBugScribe: Automatic Database Bug Reproduction from Community ReportsSuyang Zhong, Mo Sha, Sheng Wang, Fangyuan Zhou 等SIGMOD 2026 · 被引用 1 次
- SchemaRAG: A Schema-aware Retrieval-Augmented Generation Framework for Text-to-SQLDi Wu, Zetong Tang, Yi He, Xin LuoSIGMOD 2026 · 被引用 9 次
- Doctopus: Budget-aware Structural Table Extraction from Unstructured DocumentsChengliang Chai, Jiajun Li, Yuhao Deng, Yuanhao Zhong 等VLDB 2025 · 被引用 5 次
