Deep-Reporter: Deep Research for Grounded Multimodal Long-Form Generation
Fangda Ye, Kuicai Dong, Zhifei Xie, Yuxin Hu, Yihang Yin, Shurui Huang, Shikai Dong, Chen Zhang, Jianzhu Bao, Shuicheng Yan
摘要
Recent agentic search frameworks enable deep research via iterative planning and retrieval, reducing hallucinations and enhancing factual grounding. However, they remain text-centric, overlooking the multimodal evidence that characterizes real-world expert reports. We introduce a pressing task: multimodal longform generation. Accordingly, we propose DEEP-REPORTER, a unified agentic framework for grounded multimodal long-form generation. It orchestrates: (i) Agentic Multimodal Search and Filtering to retrieve and filter textual passages and information-dense visuals; (ii) Checklist-Guided Incremental Synthesis to ensure coherent image-text integration and optimal citation placement; and (iii) Recurrent Context Management to balance long-range coherence with local fluency. We develop a rigorous curation pipeline producing 8K highquality agentic traces for model optimization. We further introduce M 2 LONGBENCH, a comprehensive testbed comprising 247 research tasks across 9 domains and a stable multimodal sandbox. It enables unified multimodal assessment, fair comparison, and accessible evaluation without commercial APIs. Extensive experiments demonstrate that long-form multimodal generation is a challenging task, especially in multimodal selection and integration, and effective post-training can bridge the gap. Our code is available at https://github. com/fangda-ye/Deep-Report .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper20
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil 等ICLR 2024 · 被引用 1,798 次
- DeepResearch Bench: A Comprehensive Benchmark for Deep Research AgentsMingxuan Du, Benfeng Xu, Chiwei Zhu, Licheng Zhang 等ICLR 2026 · 被引用 250 次
- Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step QuestionsHarsh Trivedi, Niranjan Balasubramanian, Tushar Khot, Ashish SabharwalACL 2023 · 被引用 187 次
- SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated ReasoningZhenghai Xue, Longtao Zheng, Qian Liu, Yingru Li 等ICLR 2026 · 被引用 152 次
相关 Paper
- DeepImageSearch: Benchmarking Multimodal Agents for Context-Aware Image Retrieval in Visual HistoriesChenlong Deng, Mengjie Deng, Junjie Wu, Dun Zeng 等ICML 2026
- Multimodal DeepResearcher: Generating Text-Chart Interleaved Reports from Scratch with Agentic FrameworkZhaorui Yang, Bo Pan, Han Wang, Yiyao Wang 等AAAI 2026 · 被引用 12 次
- Doc-Researcher: A Unified System for Multimodal Document Parsing and Deep ResearchKuicai Dong, Shurui Huang, Fangda Ye, Wei Han 等WWW 2026 · 被引用 4 次
- Towards Knowledgeable Deep Research: Framework and BenchmarkWenxuan Liu, Zixuan Li, Long Bai, Chunmao Zhang 等SIGIR 2026
- DeepWriter: A Multi-Agent Collaboration Framework for Information-rich Ultra-long Book WritingMing Wang, Minghao Hu, Xiuli Kang, Li He 等AAAI 2026
