DocHieNet: A Large and Diverse Dataset for Document Hierarchy Parsing
Hangdi Xing, Changxu Cheng, Feiyu Gao, Zirui Shao, Zhi Yu, Jiajun Bu, Qi Zheng, Cong Yao
Abstract
Parsing documents from pixels, such as pictures and scanned PDFs, into hierarchical structures is extensively demanded in the daily routines of data storage, retrieval and understanding. However, previously the research on this topic has been largely hindered since most existing datasets are small-scale, or contain documents of only a single type, which are characterized by a lack of document diversity. Moreover, there is a significant discrepancy in the annotation standards across datasets. In this paper, we introduce a large and diverse document hierarchy parsing (DHP) dataset to compensate for the data scarcity and inconsistency problem. We aim to set a new standard as a more practical, long-standing benchmark. Meanwhile, we present a new DHP framework designed to grasp both fine-grained text content and coarsegrained pattern at layout element level, enhancing the capacity of pre-trained text-layout models in handling the multi-page and multi-level challenges in DHP. Through exhaustive experiments, we validate the effectiveness of our proposed dataset and method 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f4447cc2-31ee-4d59-beb9-962d60696e7aCited by top-tier papers4
- Is Cognition Consistent with Perception? Assessing and Mitigating Multimodal Knowledge Conflicts in Document UnderstandingZirui Shao, Feiyu Gao, Zhaoqing Zhu, Chuwei Luo et al.EMNLP 2025
- M3DocDep: Multi-modal, Multi-page, Multi-document Dependency Chunking with Large Vision-Language ModelsJoongmin Shin, Jeongbae Park, Jaehyung Seo, Heuiseok LimCVPR 2026
- HiKEY: Hierarchical Multimodal Retrieval for Open-Domain Document Question AnsweringJoongmin Shin, Gyuho Shim, Jeongbae Park, Jaehyung Seo et al.ACL 2026
- MultiDocFusion : Hierarchical and Multimodal Chunking Pipeline for Enhanced RAG on Long Industrial DocumentsJoongmin Shin, Chanjun Park, Jeongbae Park, Jaehyung Seo et al.EMNLP 2025
Builds on15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- LayoutLMv3: Pre-training for Document AI with Unified Text and Image MaskingYupan Huang, Tengchao Lv, Lei Cui, Yutong Lu et al.ACM MM 2022 · 606 citations
- LayoutLM: Pre-training of Text and Layout for Document Image UnderstandingYiheng Xu, Minghao Li, Lei Cui, Shaohan Huang et al.KDD 2020 · 575 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
Related papers
- HRDoc: Dataset and Baseline Method toward Hierarchical Reconstruction of Document StructuresJiefeng Ma, Jun Du, Pengfei Hu, Zhenrong Zhang et al.AAAI 2023 · 20 citations
- DocParser: Hierarchical Document Structure Parsing from RenderingsJohannes Rausch, Octavio Martinez, Fabian Bissig, Ce Zhang et al.AAAI 2021 · 38 citations
- OmniDocLayout: Towards Diverse Document Layout Generation via Coarse-to-Fine LLM LearningHengrui Kang, Zhuangcheng Gu, Zhiyuan Zhao, Zichen Wen et al.CVPR 2026 · 2 citations
- Graph-based Document Structure AnalysisYufan Chen, Ruiping Liu, Junwei Zheng, Di Wen et al.ICLR 2025
- MGDoc: Pre-training with Multi-granular Hierarchy for Document Image UnderstandingZilong Wang, Jiuxiang Gu, Chris Tensmeyer, Nikolaos Barmpalios et al.EMNLP 2022 · 6 citations
