DocParser: Hierarchical Document Structure Parsing from Renderings
Johannes Rausch, Octavio Martinez, Fabian Bissig, Ce Zhang, Stefan Feuerriegel
Abstract
Translating renderings (e. g. PDFs, scans) into hierarchical document structures is extensively demanded in the daily routines of many real-world applications. However, a holistic, principled approach to inferring the complete hierarchical structure of documents is missing. As a remedy, we developed "DocParser": an end-to-end system for parsing the complete document structure -including all text elements, nested figures, tables, and table cell structures. Our second contribution is to provide a dataset for evaluating hierarchical document structure parsing. Our third contribution is to propose a scalable learning framework for settings where domain-specific data are scarce, which we address by a novel approach to weak supervision that significantly improves the document structure parsing performance. Our experiments confirm the effectiveness of our proposed weak supervision: Compared to the baseline without weak supervision, it improves the mean average precision for detecting document entities by 39.1 % and improves the F1 score of classifying hierarchical relations by 35.8 %.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 79db5dfb-662f-4668-963a-d73101e6845bCited by top-tier papers10
- HRDoc: Dataset and Baseline Method toward Hierarchical Reconstruction of Document StructuresJiefeng Ma, Jun Du, Pengfei Hu, Zhenrong Zhang et al.AAAI 2023 · 20 citations
- AgentOCR: Reimagining Agent History via Optical Self-CompressionLang Feng, Fuchao Yang, Feng Chen, Xin Cheng et al.ACL 2026 · 17 citations
- SmolDocling: An Ultra-Compact Vision-Language Model for End-To-End Multi-Modal Document ConversionAhmed S. Nassar, Matteo Omenetti, Maksym Lysak, Nikolaos Livathinos et al.ICCV 2025 · 9 citations
- PosterForest: Hierarchical Multi-Agent Collaboration for Scientific Poster GenerationJiho Choi, Seojeong Park, Seongjong Song, Hyunjung ShimACL 2026 · 5 citations
- DocHieNet: A Large and Diverse Dataset for Document Hierarchy ParsingHangdi Xing, Changxu Cheng, Feiyu Gao, Zirui Shao et al.EMNLP 2024 · 3 citations
Builds on1
Related papers
- Towards Real-World Document Parsing via Realistic Scene Synthesis and Document-Aware TrainingGengluo Li, Pengyuan Lyu, Chengquan Zhang, Huawen Shen et al.CVPR 2026 · 9 citations
- Graph-based Document Structure AnalysisYufan Chen, Ruiping Liu, Junwei Zheng, Di Wen et al.ICLR 2025
- TDATR: Improving End-to-End Table Recognition via Table Detail-Aware Learning and Cell-Level Visual AlignmentChunxia Qin, Chenyu Liu, Pengcheng Xia, Jun Du et al.CVPR 2026 · 2 citations
- MGDoc: Pre-training with Multi-granular Hierarchy for Document Image UnderstandingZilong Wang, Jiuxiang Gu, Chris Tensmeyer, Nikolaos Barmpalios et al.EMNLP 2022 · 6 citations
- SciREX: A Challenge Dataset for Document-Level Information ExtractionSarthak Jain, Madeleine van Zuylen, Hannaneh Hajishirzi, Iz BeltagyACL 2020 · 9 citations
