M-LongDoc: A Benchmark For Multimodal Super-Long Document Understanding And A Retrieval-Aware Tuning Framework
Yew Ken Chia, Liying Cheng, Hou Pong Chan, Maojia Song, Chaoqun Liu, Mahani Aljunied, Soujanya Poria, Lidong Bing
摘要
The ability to understand and answer questions over documents can be useful in many business and practical applications. However, documents often contain lengthy and diverse multimodal contents such as texts, figures, and tables, which are very time-consuming for humans to read thoroughly. Hence, there is an urgent need to develop effective and automated methods to aid humans in this task. In this work, we introduce M-LongDoc, a benchmark of 851 samples, and an automated framework to evaluate the performance of large multimodal models. We further propose a retrieval-aware tuning approach for efficient and effective multimodal document reading. Compared to existing works, our benchmark consists of more recent and lengthy documents with hundreds of pages, while also requiring open-ended explanations and not just extractive answers. To our knowledge, our training framework is the first to directly address the retrieval setting for multimodal long documents. To enhance open models, we construct a training corpus in a fully automatic manner. Experiments show that our tuning approach significantly improves the correctness of model responses by 4.6%. 1 * Yew Ken and Chaoqun were students under the Joint PhD Program between Alibaba and their corresponding university. Work done while Liying, Mahani, and Lidong were at Alibaba. † Corresponding authors. 1 Our multimodal benchmark, training corpus, and source code are publicly available at https://multimodal-documents.github.io/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Scaling Language-centric Omnimodal Representation LearningChenghao Xiao, Hou Pong Chan, Hao Zhang, Weiwen Xu 等NeurIPS 2025 · 被引用 25 次
- Resolving Evidence Sparsity: Agentic Context Engineering for Long-Document UnderstandingKeliang Liu, Zizhi Chen, Mingcheng Li, Jingqun Tang 等CVPR 2026 · 被引用 19 次
- Deep-Reporter: Deep Research for Grounded Multimodal Long-Form GenerationFangda Ye, Kuicai Dong, Zhifei Xie, Yuxin Hu 等ACL 2026 · 被引用 2 次
- Untied Ulysses: Memory-Efficient Context Parallelism via Headwise ChunkingRavi Ghadia, Maksim Abraham, Sergei Vorobyov, Max RyabininICML 2026
- Understanding the Behaviors of Environment-aware Information RetrievalRuifeng Yuan, Chaohao Yuan, David Dai, Yu Rong 等ACL 2026
它引用的顶会 Paper12
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Large Language Models Can Be Easily Distracted by Irrelevant ContextFreda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales 等ICML 2023 · 被引用 970 次
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang 等EMNLP 2023 · 被引用 549 次
相关 Paper
- Benchmarking Retrieval-Augmented Generation in Multi-Modal ContextsZhenghao Liu, Xingsheng Zhu, Tianshuo Zhou, Xinyi Zhang 等ACM MM 2025 · 被引用 4 次
- M4LE: A Multi-Ability Multi-Range Multi-Task Multi-Domain Long-Context Evaluation Benchmark for Large Language ModelsWai-Chung Kwan, Xingshan Zeng, Yufei Wang, Yusen Sun 等ACL 2024 · 被引用 3 次
- MMDocIR: Benchmarking Multimodal Retrieval for Long DocumentsKuicai Dong, Yujing Chang, Derrick-Goh-Xin Deik, Dexun Li 等EMNLP 2025 · 被引用 1 次
- LongBench: A Bilingual, Multitask Benchmark for Long Context UnderstandingYushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu 等ACL 2024 · 被引用 94 次
- Leave No Document Behind: Benchmarking Long-Context LLMs with Extended Multi-Doc QAMinzheng Wang, Longze Chen, Cheng Fu, Shengyi Liao 等EMNLP 2024 · 被引用 9 次
