HybriDLA: Hybrid Generation for Document Layout Analysis
Yufan Chen, Omar Moured, Ruiping Liu, Junwei Zheng, Kunyu Peng, Jiaming Zhang, Rainer Stiefelhagen
Abstract
Conventional document layout analysis (DLA) traditionally depends on empirical priors or a fixed set of learnable queries executed in a single forward pass. While sufficient for early-generation documents with a small, predetermined number of regions, this paradigm struggles with contemporary documents, which exhibit diverse element counts and increasingly complex layouts. To address challenges posed by modern documents, we present HybriDLA, a novel generative framework that unifies diffusion and autoregressive decoding within a single layer. The diffusion component iteratively refines bounding-box hypotheses, whereas the autoregressive component injects semantic and contextual awareness, enabling precise region prediction even in highly varied layouts. To further enhance detection quality, we design a multi-scale feature-fusion encoder that captures both fine-grained and high-level visual cues. This architecture elevates performance to 83.5% mean Average Precision (mAP). Extensive experiments on the DocLayNet and M6Doc benchmarks demonstrate that HybriDLA sets a state-of-the-art performance, outperforming previous approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7d6c9655-0fa1-4e91-b636-600909fc7746Builds on20
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object DetectionHao Zhang, Feng Li, Shilong Liu, Lei Zhang et al.ICLR 2023 · 753 citations
- DiffusionDet: Diffusion Model for Object DetectionShoufa Chen, Peize Sun, Yibing Song, Ping LuoICCV 2023 · 715 citations
- DETRs with Collaborative Hybrid Assignments TrainingZhuofan Zong, Guanglu Song, Yu LiuICCV 2023 · 594 citations
Related papers
- M2Doc: A Multi-Modal Fusion Approach for Document Layout AnalysisNing Zhang, Hiuyi Cheng, Jiayu Chen, Zongyuan Jiang et al.AAAI 2024 · 16 citations
- LLM-Guided Probabilistic Fusion for Label-Efficient Document Layout AnalysisIbne Farabi Shihab, Sanjeda Akter, Anuj SharmaCVPR 2026
- Infinite-Precision Autoregressive Modeling for Vector Graphics and LayoutsYeonsang Shin, Insoo Kim, Bongkeun Kim, Keonwoo Bae et al.ICML 2026
- FastHybrid: Accelerating Hybrid Autoregressive Image Generation with Lookahead and Guided DecodingZhengguo Jiang, Fang Zhang, YongXiang Hua, Bocheng Li et al.CVPR 2026
- Visual Prototype Conditioned Focal Region Generation for UAV-Based Object DetectionWenhao Li, Zimeng Wu, Yu Wu, Zehua Fu et al.CVPR 2026
