MultiDocFusion : Hierarchical and Multimodal Chunking Pipeline for Enhanced RAG on Long Industrial Documents
Joongmin Shin, Chanjun Park, Jeongbae Park, Jaehyung Seo, Heuiseok Lim
摘要
RAG-based QA has emerged as a powerful method for processing long industrial documents. However, conventional text chunking approaches often neglect the complex structures of long industrial documents, causing information loss and reduced answer quality. To address this, we introduce MultiDocFusion, a multimodal chunking pipeline that integrates: (i) detection of document regions using visionbased document parsing, (ii) text extraction from these regions via OCR, (iii) reconstruction of document structure into a hierarchical tree using large language model (LLM)based document section hierarchical parsing (DSHP-LLM), and (iv) construction of hierarchical chunks through DFS-based Grouping. Extensive experiments across industrial benchmarks demonstrate that MultiDocFusion improves retrieval precision by 8-15% and ANLS QA scores by 2-3% compared to baselines, emphasizing the critical role of explicitly leveraging document hierarchy for multimodal document-based QA. These significant performance gains underscore the necessity of structure-aware chunking in enhancing the fidelity of RAG-based QA systems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- M3DocDep: Multi-modal, Multi-page, Multi-document Dependency Chunking with Large Vision-Language ModelsJoongmin Shin, Jeongbae Park, Jaehyung Seo, Heuiseok LimCVPR 2026
- HiKEY: Hierarchical Multimodal Retrieval for Open-Domain Document Question AnsweringJoongmin Shin, Gyuho Shim, Jeongbae Park, Jaehyung Seo 等ACL 2026
它引用的顶会 Paper12
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- TrOCR: Transformer-Based Optical Character Recognition with Pre-trained ModelsMinghao Li, Tengchao Lv, Jingye Chen, Lei Cui 等AAAI 2023 · 被引用 607 次
- Scene Text Visual Question AnsweringAli Furkan Biten, Rubèn Tito, Andrés Mafla, Lluís Gómez i Bigorda 等ICCV 2019 · 被引用 482 次
相关 Paper
- MoDora: Tree-Based Semi-Structured Document Analysis SystemBangrui Xu, Qihang Yao, Zirui Tang, Xuanhe Zhou 等SIGMOD 2026
- RegionRAG: Region-level Retrieval-Augmented Generation for Visual Document UnderstandingYinglu Li, Zhiying Lu, Zhihang Liu, Yiwei Sun 等AAAI 2026 · 被引用 2 次
- Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document UnderstandingSensen Gao, Shanshan Zhao, Xu Jiang, Lunhao Duan 等ACL 2026 · 被引用 7 次
- VDocRAG: Retrieval-Augmented Generation over Visually-Rich DocumentsRyota Tanaka, Taichi Iki, Taku Hasegawa, Kyosuke Nishida 等CVPR 2025
- Beyond Chunking: Discourse-Aware Hierarchical Retrieval for Long Document Question AnsweringHuiyao Chen, Yi Yang, Yinghui Li, Meishan Zhang 等ACL 2026 · 被引用 6 次
