DocAgent: An Agentic Framework for Multi-Modal Long-Context Document Understanding
Li Sun, Liu He, Shuyue Jia, Yangfan He, Chenyu You
Abstract
Recent advances in large language models (LLMs) have demonstrated significant promise in document understanding and questionanswering. Despite the progress, existing approaches can only process short documents due to limited context length or fail to fully leverage multi-modal information. In this work, we introduce DocAgent, a multi-agent framework for long-context document understanding that imitates the human reading practice. Specifically, we first extract a structured, tree-formatted outline from documents to help agents identify relevant sections efficiently. Further, we develop an interactive reading interface that enables agents to query and retrieve various types of content dynamically. To ensure answer reliability, we introduce a reviewer agent that cross-checks responses using complementary sources and maintains a task-agnostic memory bank to facilitate knowledge sharing across tasks. We evaluate our method on two long-context document understanding benchmarks, where it bridges the gap to human-level performance by surpassing competitive baselines, while maintaining a short context length. Our code is available at https://github.com/lisun-ai/DocAgent .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aacca0f5-c656-4c83-97f2-322dd9d27bf0Cited by top-tier papers2
- When to Think, When to Speak: Learning Disclosure Policies for LLM ReasoningJiaqi Wei, Xuehang Guo, Pengfei Yu, Xiang Zhang et al.ICML 2026 · 2 citations
- MoDora: Tree-Based Semi-Structured Document Analysis SystemBangrui Xu, Qihang Yao, Zirui Tang, Xuanhe Zhou et al.SIGMOD 2026
Builds on11
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Large Language Models Can Be Easily Distracted by Irrelevant ContextFreda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales et al.ICML 2023 · 970 citations
- LayoutLMv3: Pre-training for Document AI with Unified Text and Image MaskingYupan Huang, Tengchao Lv, Lei Cui, Yutong Lu et al.ACM MM 2022 · 606 citations
- LayoutLM: Pre-training of Text and Layout for Document Image UnderstandingYiheng Xu, Minghao Li, Lei Cui, Shaohan Huang et al.KDD 2020 · 575 citations
- Language Agent Tree Search Unifies Reasoning, Acting, and Planning in Language ModelsAndy Zhou, Kai Yan, Michal Shlapentokh-Rothman, Haohan Wang et al.ICML 2024 · 443 citations
Related papers
- A Human-Inspired Reading Agent with Gist Memory of Very Long ContextsKuang-Huei Lee, Xinyun Chen, Hiroki Furuta, John F. Canny et al.ICML 2024 · 106 citations
- Chain of Agents: Large Language Models Collaborating on Long-Context TasksYusen Zhang, Ruoxi Sun, Yanfei Chen, Tomas Pfister et al.NeurIPS 2024 · 297 citations
- LVAgent: Long Video Understanding by Multi-Round Dynamical Collaboration of MLLM AgentsBoyu Chen, Zhengrong Yue, Siran Chen, Zikang Wang et al.ICCV 2025 · 12 citations
- SlideAgent: Hierarchical Agentic Framework for Multi-Page Visual Document UnderstandingYiqiao Jin, Rachneet Kaur, Zhen Zeng, Sumitra Ganesh et al.ACL 2026 · 1 citation
- Learning to Summarize by Learning to Quiz: Adversarial Agentic Collaboration for Long Document SummarizationWeixuan Wang, Minghao Wu, Barry Haddow, Alexandra BirchICLR 2026 · 3 citations
