DIG: Complex Layout Document Image Generation with Authentic-looking Text for Enhancing Layout Analysis
Dehao Ying, Fengchang Yu, Haihua Chen, Wei Lu
摘要
Even though significant progress has been made in standardizing document layout analysis, complex layout documents like magazines and newspapers still present challenges. Models trained on standardized documents struggle with these complexities, and the high cost of annotating such documents limits dataset availability. To address this, we propose the Complex Layout Document Image Generation (DIG) model, which can generate diverse document images with complex layouts and authentic-looking text, aiding in layout analysis model training. Concretely, we first pre-train DIG on a large-scale document dataset with a text-sensitive loss function to address the issue of unreal generation of text regions. Then, we fine-tune it with a small number of documents with complex layouts to generate new images with the same layout. Additionally, we use a layout generation model to create new layouts, enhancing data diversity. Finally, we design a box-wise quality scoring function to filter out low-quality regions during layout analysis model training to enhance the effectiveness of using the generated images. Experimental results on the DSSE-200 and PRImA datasets show when incorporating generated images from DIG, the mAP of the layout analysis model is improved from 47.05 to 56.07 and from 53.80 to 62.26, respectively, which is a 19.17% and 15.72% enhancement compared to the baseline.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- OmniDocLayout: Towards Diverse Document Layout Generation via Coarse-to-Fine LLM LearningHengrui Kang, Zhuangcheng Gu, Zhiyuan Zhao, Zichen Wen 等CVPR 2026 · 被引用 2 次
- Graph-based Document Structure AnalysisYufan Chen, Ruiping Liu, Junwei Zheng, Di Wen 等ICLR 2025
- RoDLA: Benchmarking the Robustness of Document Layout Analysis ModelsYufan Chen, Jiaming Zhang, Kunyu Peng, Junwei Zheng 等CVPR 2024
- DocHieNet: A Large and Diverse Dataset for Document Hierarchy ParsingHangdi Xing, Changxu Cheng, Feiyu Gao, Zirui Shao 等EMNLP 2024 · 被引用 3 次
- Vision Grid Transformer for Document Layout AnalysisCheng Da, Chuwei Luo, Qi Zheng, Cong YaoICCV 2023 · 被引用 63 次
