AnyDoc: Enhancing Document Generation via Large-Scale HTML/CSS Data Synthesis and Height-Aware Reinforcement Optimization
Jiawei Lin, Wanrong Zhu, Vlad I Morariu, Christopher Tensmeyer
Abstract
Document generation has emerged as a crucial task for automating the creation of visually appealing and well-structured content across diverse domains. Existing methods in this field, however, suffer from some limitations in terms of application scope, document representation and dataset coverage, which greatly restricts the capabilities of document generation models. To address these challenges, we propose OmniDoc, a framework that introduces HTML/CSS as a novel document representation given its inherent advantages in hierarchical structure modeling. Leveraging HTML/CSS, OmniDoc establishes a scalable data synthesis pipeline to curate DocHTML, a large-scale document dataset containing 265,206 high-quality samples. Each document in DocHTML includes complete metadata annotations, structured HTML/CSS source code, synthesized visual assets, and rendered screenshots, spanning diverse categories, styles, and complexity levels to ensure comprehensive coverage. OmniDoc then utilizes DocHTML to fine-tune the multimodal large language models, empowering them remarkable document generation capabilities on three practical tasks: intention-to-document, document derendering, and element-to-document. To address the content overflow issues found in the fine-tuned models, we incorporate a height-aware post-training method within OmniDoc based on Group Relative Policy Optimization. By carefully designing the reward function to measure the alignment between predicted and target document heights, OmniDoc effectively alleviates the overflow problem, further enhancing model performance. Qualitative and quantitative results demonstrate the superiority of OmniDoc over baseline models across all three tasks. Extensive ablation studies manifest the effectiveness of the HTML/CSS representation, curated dataset, and height-aware reinforcement optimization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on21
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- CanvasVAE: Learning to Generate Vector Graphic DocumentsKota YamaguchiICCV 2021 · 103 citations
Related papers
- OmniDocLayout: Towards Diverse Document Layout Generation via Coarse-to-Fine LLM LearningHengrui Kang, Zhuangcheng Gu, Zhiyuan Zhao, Zichen Wen et al.CVPR 2026 · 2 citations
- Adaptive Markup Language Generation for Contextually-Grounded Visual Document UnderstandingHan Xiao, Yina Xie, Guanxin Tan, Yinghao Chen et al.CVPR 2025
- Omni-Mol: Multitask Molecular Model for Any-to-any ModalitiesChengxin Hu, Hao Li, Yihe Yuan, Zezheng Song et al.NeurIPS 2025 · 5 citations
- Relation-Rich Visual Document Generator for Visual Information ExtractionZi-Han Jiang, Chien-Wei Lin, Wei-Hua Li, Hsuan-Tung Liu et al.CVPR 2025
- Towards Real-World Document Parsing via Realistic Scene Synthesis and Document-Aware TrainingGengluo Li, Pengyuan Lyu, Chengquan Zhang, Huawen Shen et al.CVPR 2026 · 9 citations
