MultiDocFusion : Hierarchical and Multimodal Chunking Pipeline for Enhanced RAG on Long Industrial Documents
Joongmin Shin, Chanjun Park, Jeongbae Park, Jaehyung Seo, Heuiseok Lim
Abstract
RAG-based QA has emerged as a powerful method for processing long industrial documents. However, conventional text chunking approaches often neglect the complex structures of long industrial documents, causing information loss and reduced answer quality. To address this, we introduce MultiDocFusion, a multimodal chunking pipeline that integrates: (i) detection of document regions using visionbased document parsing, (ii) text extraction from these regions via OCR, (iii) reconstruction of document structure into a hierarchical tree using large language model (LLM)based document section hierarchical parsing (DSHP-LLM), and (iv) construction of hierarchical chunks through DFS-based Grouping. Extensive experiments across industrial benchmarks demonstrate that MultiDocFusion improves retrieval precision by 8-15% and ANLS QA scores by 2-3% compared to baselines, emphasizing the critical role of explicitly leveraging document hierarchy for multimodal document-based QA. These significant performance gains underscore the necessity of structure-aware chunking in enhancing the fidelity of RAG-based QA systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c4c5d09e-cbbf-4034-b472-14777beff2dcCited by top-tier papers2
- M3DocDep: Multi-modal, Multi-page, Multi-document Dependency Chunking with Large Vision-Language ModelsJoongmin Shin, Jeongbae Park, Jaehyung Seo, Heuiseok LimCVPR 2026
- HiKEY: Hierarchical Multimodal Retrieval for Open-Domain Document Question AnsweringJoongmin Shin, Gyuho Shim, Jeongbae Park, Jaehyung Seo et al.ACL 2026
Builds on12
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- TrOCR: Transformer-Based Optical Character Recognition with Pre-trained ModelsMinghao Li, Tengchao Lv, Jingye Chen, Lei Cui et al.AAAI 2023 · 607 citations
- Scene Text Visual Question AnsweringAli Furkan Biten, Rubèn Tito, Andrés Mafla, Lluís Gómez i Bigorda et al.ICCV 2019 · 482 citations
Related papers
- MoDora: Tree-Based Semi-Structured Document Analysis SystemBangrui Xu, Qihang Yao, Zirui Tang, Xuanhe Zhou et al.SIGMOD 2026
- RegionRAG: Region-level Retrieval-Augmented Generation for Visual Document UnderstandingYinglu Li, Zhiying Lu, Zhihang Liu, Yiwei Sun et al.AAAI 2026 · 2 citations
- Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document UnderstandingSensen Gao, Shanshan Zhao, Xu Jiang, Lunhao Duan et al.ACL 2026 · 7 citations
- VDocRAG: Retrieval-Augmented Generation over Visually-Rich DocumentsRyota Tanaka, Taichi Iki, Taku Hasegawa, Kyosuke Nishida et al.CVPR 2025
- Beyond Chunking: Discourse-Aware Hierarchical Retrieval for Long Document Question AnsweringHuiyao Chen, Yi Yang, Yinghui Li, Meishan Zhang et al.ACL 2026 · 6 citations
