Doc-Researcher: A Unified System for Multimodal Document Parsing and Deep Research
Kuicai Dong, Shurui Huang, Fangda Ye, Wei Han, Zhi Zhang, Dexun Li, Wenjun Li, Qu Yang, Gang Wang, Yichao Wang, Chen Zhang, Yong Liu
Abstract
Deep Research systems have revolutionized how LLMs solve complex questions through iterative reasoning and evidence gathering. However, current systems remain fundamentally constrained to textual web data, overlooking the vast knowledge embedded in multimodal documents: scientific papers, technical reports, and financial documents where critical information exists in figures, tables, charts, and equations. Processing such documents demands sophisticated parsing to preserve visual semantics, intelligent chunking to maintain structural coherence, and adaptive retrieval across modalities, which are capabilities absent in existing systems. In response, we present Doc-Researcher, a unified system that bridges this gap through three integrated components: (i) deep multimodal parsing that preserves layout structure and visual semantics while creating multi-granular representations from chunk to document level, (ii) systematic retrieval architecture supporting text-only, vision-only, and hybrid paradigms with dynamic granularity selection, and (iii) iterative multi-agent workflows that decompose complex queries, progressively accumulate evidence, and synthesize comprehensive answers across documents and modalities. To enable rigorous evaluation, we introduce M4DocBench, the first benchmark for Multi-modal, Multi-hop, Multi-document, and Multi-turn deep research. Featuring 158 expert-annotated questions with complete evidence chains across 304 documents, M4DocBench tests capabilities that existing benchmarks cannot assess. Experiments demonstrate that Doc-Researcher achieves 50.6% accuracy, 3.4× better than state-of-the-art baselines, validating that effective document research requires not just better retrieval, but fundamentally deep parsing that preserve multimodal integrity and support iterative research. Our work establishes a new paradigm for conducting deep research on multimodal document collections. CCS Concepts • Information systems → Multimodal Deep Research.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2ca55095-3149-4931-a9db-9a5adaf62c56Cited by top-tier papers2
- When Hard Negatives Hurt: Bridging the Generative Discriminative Gap in Hard Negative Synthesis for RetrievalZhicheng Zhang, Jiwei Tang, Kuicai Dong, Xiaopeng Li et al.KDD 2026 · 1 citation
- Learning More from Less: Exploiting Counterfactuals for Data-Efficient Chart UnderstandingJianzhu Bao, Haozhen Zhang, Kuicai Dong, Bozhi Wu et al.ACL 2026
Builds on13
- WebThinker: Empowering Large Reasoning Models with Deep Research CapabilityXiaoxi Li, Jiajie Jin, Guanting Dong, Hongjin Qian et al.NeurIPS 2025 · 354 citations
- DeepResearch Bench: A Comprehensive Benchmark for Deep Research AgentsMingxuan Du, Benfeng Xu, Chiwei Zhu, Licheng Zhang et al.ICLR 2026 · 250 citations
- Document Understanding Dataset and Evaluation (DUDE)Jordy Van Landeghem, Rafal Powalski, Rubèn Tito, Dawid Jurkiewicz et al.ICCV 2023 · 130 citations
- WebSailor-V2: Bridging the Chasm to Proprietary Agents via Synthetic Data and Scalable Reinforcement LearningKuan Li, Zhongwang Zhang, Huifeng Yin, Rui Ye et al.ICLR 2026 · 65 citations
- WebWeaver: Structuring Web-Scale Evidence with Dynamic Outlines for Open-Ended Deep ResearchZijian Li, Xin Guan, Bo Zhang, Shen Huang et al.ICLR 2026 · 48 citations
Related papers
- DocRetriever: A Plug-and-Play Framework for Multimodal Document Retrieval with Comprehensive BenchmarkRuofan Hu, Menghui Zhu, Jieming Zhu, Bo Chen et al.KDD 2026 · 1 citation
- WebWatcher: Breaking New Frontiers of Vision-Language Deep Research AgentXinyu Geng, Peng Xia, Zhen Zhang, Xinyu Wang et al.ICLR 2026 · 79 citations
- Towards Knowledgeable Deep Research: Framework and BenchmarkWenxuan Liu, Zixuan Li, Long Bai, Chunmao Zhang et al.SIGIR 2026
- MMDocIR: Benchmarking Multimodal Retrieval for Long DocumentsKuicai Dong, Yujing Chang, Derrick-Goh-Xin Deik, Dexun Li et al.EMNLP 2025 · 1 citation
- Multimodal DeepResearcher: Generating Text-Chart Interleaved Reports from Scratch with Agentic FrameworkZhaorui Yang, Bo Pan, Han Wang, Yiyao Wang et al.AAAI 2026 · 12 citations
