M3DocDep: Multi-modal, Multi-page, Multi-document Dependency Chunking with Large Vision-Language Models
Joongmin Shin, Jeongbae Park, Jaehyung Seo, Heuiseok Lim
Abstract
In long, multi-page industrial documents, retrieval-augmented generation (RAG) depends heavily on whether chunk boundaries follow the document's true structure. Existing text-centric chunkers and generative hierarchy parsers often miss cross-page parent-child relations, figure/table-caption bindings, and boundary cues, which leads to fragmented or redundant chunks and degrades both retrieval and answer quality. We propose M3DocDep, an LVLM-based pipeline that first recovers block-level dependencies and then constructs chunks along the recovered document tree. The pipeline uses SharedDet as a common DP+OCR preprocessing layer, extracts multimodal block embeddings with boundary-aware SoftROI pooling, scores candidate parent-child edges with a biaffine head, decodes a globally valid dependency tree with MST constraints, and builds tree-guided chunks annotated with section paths and page ranges. Under a shared-block evaluation protocol, M3DocDep improves STEDS by +28.5 to +39.6 percent on DHP benchmarks, retrieval nDCG by +1.1 to +15.3 percent, and QA ANLS by +4.5 to +15.3 percent on corpus-level RAG benchmarks. These results show that recovering document dependencies before chunking yields more coherent retrieval units for long, multi-page multimodal documents.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 778736ed-d849-4baf-b45e-e85e706c6ac5Cited by top-tier papers1
Ask how each one uses itBuilds on11
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Scene Text Visual Question AnsweringAli Furkan Biten, Rubèn Tito, Andrés Mafla, Lluís Gómez i Bigorda et al.ICCV 2019 · 482 citations
- Document Understanding Dataset and Evaluation (DUDE)Jordy Van Landeghem, Rafal Powalski, Rubèn Tito, Dawid Jurkiewicz et al.ICCV 2023 · 130 citations
- Recurrent Chunking Mechanisms for Long-Text Machine Reading ComprehensionHongyu Gong, Yelong Shen, Dian Yu, Jianshu Chen et al.ACL 2020 · 39 citations
Related papers
- MultiDocFusion : Hierarchical and Multimodal Chunking Pipeline for Enhanced RAG on Long Industrial DocumentsJoongmin Shin, Chanjun Park, Jeongbae Park, Jaehyung Seo et al.EMNLP 2025
- LAD-RAG: Layout-aware Dynamic RAG for Visually-Rich Document UnderstandingZhivar Sourati, Zheng Wang, Marianne Menglin Liu, Yazhe Hu et al.ACL 2026 · 5 citations
- SmartChunk Retrieval: Query-Aware Chunk Compression with Planning for Efficient Document RAGXuechen Zhang, Koustava Goswami, Samet Oymak, Jiasi Chen et al.ICLR 2026 · 1 citation
- TH-RAG : Topic-Based Hierarchical Knowledge Graphs for Robust Multi-hop Reasoning in Graph-based RAG SystemsJungHyoun Kim, Soohyeong Kim, Seok Jun Hwang, Jeonghyeon Park et al.ACL 2026
- RegionRAG: Region-level Retrieval-Augmented Generation for Visual Document UnderstandingYinglu Li, Zhiying Lu, Zhihang Liu, Yiwei Sun et al.AAAI 2026 · 2 citations
