M6Doc: A Large-Scale Multi-Format, Multi-Type, Multi-Layout, Multi-Language, Multi-Annotation Category Dataset for Modern Document Layout Analysis
Hiuyi Cheng, Peirong Zhang, Sihang Wu, Jiaxin Zhang, Qiyuan Zhu, Zecheng Xie, Jing Li, Kai Ding, Lianwen Jin
Abstract
Document layout analysis is a crucial prerequisite for document understanding, including document retrieval and conversion. Most public datasets currently contain only PDF documents and lack realistic documents. Models trained on these datasets may not generalize well to real-world scenarios. Therefore, this paper introduces a large and diverse document layout analysis dataset called M 6 Doc. The M 6 designation represents six properties: (1) Multi-Format (including scanned, photographed, and PDF documents); (2) Multi-Type (such as scientific articles, textbooks, books, test papers, magazines, newspapers, and notes); (3) Multi-Layout (rectangular, Manhattan, non-Manhattan, and multi-column Manhattan); (4) Multi-Language (Chinese and English); (5) Multi-Annotation Category (74 types of annotation labels with 237,116 annotation instances in 9,080 manually annotated pages); and (6) Modern documents. Additionally, we propose a transformer-based document layout analysis method called TransDLANet, which leverages an adaptive element matching mechanism that enables query embedding to better match ground truth to improve recall, and constructs a segmentation branch for more precise document image instance segmentation. We conduct a comprehensive evaluation of M 6 Doc with various layout analysis methods and demonstrate its effectiveness. TransDLANet achieves stateof-the-art performance on M 6 Doc with 64.5% mAP. The M 6 Doc dataset will be available at https://github. com/HCIILAB/M6Doc .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers20
- Spot the Fake: Large Multimodal Model-Based Synthetic Image Detection with Artifact ExplanationSiwei Wen, Junyan Ye, Peilin Feng, Hengrui Kang et al.NeurIPS 2025 · 82 citations
- Detect Anything via Next Point PredictionQing Jiang, Junan Huo, Xingyu Chen, Yuda Xiong et al.CVPR 2026 · 79 citations
- LayoutLLM: Layout Instruction Tuning with Large Language Models for Document UnderstandingChuwei Luo, Yufan Shen, Zhaoqing Zhu, Qi Zheng et al.CVPR 2024 · 39 citations
- DocKylin: A Large Multimodal Model for Visual Document Understanding with Efficient Visual SlimmingJiaxin Zhang, Wentao Yang, Songxuan Lai, Zecheng Xie et al.AAAI 2025 · 39 citations
- M2Doc: A Multi-Modal Fusion Approach for Document Layout AnalysisNing Zhang, Hiuyi Cheng, Jiayu Chen, Zongyuan Jiang et al.AAAI 2024 · 16 citations
Builds on9
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
- Generalized Focal Loss: Learning Qualified and Distributed Bounding Boxes for Dense Object DetectionXiang Li, Wenhai Wang, Lijun Wu, Shuo Chen et al.NeurIPS 2020 · 2,118 citations
- SOLOv2: Dynamic and Fast Instance SegmentationXinlong Wang, Rufeng Zhang, Tao Kong, Lei Li et al.NeurIPS 2020 · 1,193 citations
- Instances as QueriesYuxin Fang, Shusheng Yang, Xinggang Wang, Yu Li et al.ICCV 2021 · 331 citations
Related papers
- OmniDocLayout: Towards Diverse Document Layout Generation via Coarse-to-Fine LLM LearningHengrui Kang, Zhuangcheng Gu, Zhiyuan Zhao, Zichen Wen et al.CVPR 2026 · 2 citations
- Vision Grid Transformer for Document Layout AnalysisCheng Da, Chuwei Luo, Qi Zheng, Cong YaoICCV 2023 · 63 citations
- Graph-based Document Structure AnalysisYufan Chen, Ruiping Liu, Junwei Zheng, Di Wen et al.ICLR 2025
- DocSAM: Unified Document Image Segmentation via Query Decomposition and Heterogeneous Mixed LearningXiao-Hui Li, Fei Yin, Cheng-Lin LiuCVPR 2025
- HybriDLA: Hybrid Generation for Document Layout AnalysisYufan Chen, Omar Moured, Ruiping Liu, Junwei Zheng et al.AAAI 2026
