LaTeX2Layout: High-Fidelity, Scalable Document Layout Annotation Pipeline for Layout Detection
Feijiang Han, Zelong Wang, Bowen Wang, Xinxin Liu, Skyler Cheung, Delip Rao, Chris Callison-Burch, Lyle H. Ungar
Abstract
General-purpose Vision-Language Models (VLMs) are increasingly integral to modern AI systems for document understanding, yet their ability to perform fine-grained layout analysis remains severely underdeveloped. Overcoming this limitation requires large-scale, high-fidelity training datasets. However, current annotation methods that rely on parsing rendered PDFs are costly, error-prone, and difficult to scale. We propose a different paradigm: extracting ground-truth layout directly from the L A T E X compilation process rather than the final PDF. We present LaTeX2Layout, a generalizable procedural pipeline that recovers pixel-accurate bounding boxes and reading order from compiler traces. This enables the generation of a 140K-page dataset, including 120K programmatically generated synthetic variants that more than double the layout diversity of real-world data. Using this dataset, we fine-tune an efficient 3B-parameter VLM with an easy-to-hard curriculum that accelerates convergence. Our model achieves Kendall's ω = 0.95 for reading order and mAP@50= 0.91 for element grounding, delivering nearly 200% relative improvement over strong zero-shot baselines such as GPT-4o and Claude-3.7.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 52b5e5f0-efb1-4b98-baed-c745d3bd8fafCited by top-tier papers3
- ZeroTuning: Unlocking the Initial Token's Power to Enhance Large Language Models Without TrainingFeijiang Han, Xiaodong Yu, Jianheng Tang, Delip Rao et al.ICLR 2026 · 17 citations
- Evoking User Memory: Personalizing LLM via Recollection-Familiarity Adaptive RetrievalYingyi Zhang, Junyi Li, Wenlin Zhang, Pengyue Jia et al.ICLR 2026 · 11 citations
- Exposing and Defending the Achilles' Heel of Video Mixture-of-ExpertsSongping Wang, Qinglong Liu, Yueming Lyu, Ning Li et al.ICLR 2026 · 3 citations
Builds on7
- LayoutLMv3: Pre-training for Document AI with Unified Text and Image MaskingYupan Huang, Tengchao Lv, Lei Cui, Yutong Lu et al.ACM MM 2022 · 606 citations
- LayoutLM: Pre-training of Text and Layout for Document Image UnderstandingYiheng Xu, Minghao Li, Lei Cui, Shaohan Huang et al.KDD 2020 · 575 citations
- LiLT: A Simple yet Effective Language-Independent Layout Transformer for Structured Document UnderstandingJiapeng Wang, Lianwen Jin, Kai DingACL 2022 · 188 citations
- Vision Grid Transformer for Document Layout AnalysisCheng Da, Chuwei Luo, Qi Zheng, Cong YaoICCV 2023 · 63 citations
- Relation-Rich Visual Document Generator for Visual Information ExtractionZi-Han Jiang, Chien-Wei Lin, Wei-Hua Li, Hsuan-Tung Liu et al.CVPR 2025
Related papers
- OmniDocLayout: Towards Diverse Document Layout Generation via Coarse-to-Fine LLM LearningHengrui Kang, Zhuangcheng Gu, Zhiyuan Zhao, Zichen Wen et al.CVPR 2026 · 2 citations
- OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive AnnotationsLinke Ouyang, Yuan Qu, Hongbin Zhou, Jiawei Zhu et al.CVPR 2025
- TexOCR: Advancing Document OCR Models for Compilable Page-to-LaTeX ReconstructionChengye Wang, Lin Fu, Zexi Kuang, Yilun ZhaoACL 2026
- Adaptive Markup Language Generation for Contextually-Grounded Visual Document UnderstandingHan Xiao, Yina Xie, Guanxin Tan, Yinghao Chen et al.CVPR 2025
- Towards Real-World Document Parsing via Realistic Scene Synthesis and Document-Aware TrainingGengluo Li, Pengyuan Lyu, Chengquan Zhang, Huawen Shen et al.CVPR 2026 · 9 citations
