ALDEN: Reinforcement Learning for Active Navigation and Evidence Gathering in Long Documents
Tianyu Yang, Terry Ruas, Yijun Tian, Jan Philip Wahle, Daniel Kurzawe, Bela Gipp
摘要
Vision-language models (VLMs) excel at interpreting text-rich images but struggle with long, visually complex documents that demand analysis and integration of information spread across multiple pages. Existing approaches typically rely on fixed reasoning templates or rigid pipelines, which force VLMs into a passive role and hinder both efficiency and generalization. We present Active Long-DocumEnt Navigation (ALDEN), a multi-turn reinforcement learning framework that fine-tunes VLMs as interactive agents capable of actively navigating long, visually rich documents. ALDEN introduces a novel fetch action that directly accesses the page by index, complementing the classic search action and better exploiting document structure. For dense process supervision and efficient training, we propose a rule-based cross-level reward that provides both turn- and token-level signals. To address the empirically observed training instability caused by numerous visual tokens from long documents, we further propose a visual-semantic anchoring mechanism that applies a dual-path KL-divergence constraint to stabilize visual and textual representations separately during training. Trained on a corpus constructed from three open-source datasets, ALDEN achieves state-of-the-art performance on five long-document benchmarks. Overall, ALDEN marks a step beyond passive document reading toward agents that autonomously navigate and reason across long, visually rich documents, offering a robust path to more accurate and efficient long-document understanding.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- SlideVQA: A Dataset for Document Visual Question Answering on Multiple ImagesRyota Tanaka, Kyosuke Nishida, Kosuke Nishida, Taku Hasegawa 等AAAI 2023 · 被引用 178 次
- ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RLYifei Zhou, Andrea Zanette, Jiayi Pan, Sergey Levine 等ICML 2024 · 被引用 163 次
- Document Understanding Dataset and Evaluation (DUDE)Jordy Van Landeghem, Rafal Powalski, Rubèn Tito, Dawid Jurkiewicz 等ICCV 2023 · 被引用 130 次
- V-Doc : Visual questions answers with DocumentsYihao Ding, Zhe Huang, Runlin Wang, Yanhang Zhang 等CVPR 2022 · 被引用 19 次
相关 Paper
- Resolving Evidence Sparsity: Agentic Context Engineering for Long-Document UnderstandingKeliang Liu, Zizhi Chen, Mingcheng Li, Jingqun Tang 等CVPR 2026 · 被引用 19 次
- SV-RAG: LoRA-Contextualizing Adaptation of MLLMs for Long Document UnderstandingJian Chen, Ruiyi Zhang, Yufan Zhou, Tong Yu 等ICLR 2025
- DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document UnderstandingHao Yan, Yuliang Liu, Xingchen Liu, Yuyi Zhang 等CVPR 2026 · 被引用 9 次
- DocLens: A Tool-Augmented Multi-Agent Framework for Long Visual Document UnderstandingDawei Zhu, Rui Meng, Jiefeng Chen, Sujian Li 等ACL 2026 · 被引用 10 次
- UniDocVLM: Enhancing Visual Reasoning and Document Understanding for VLM via Reinforcement LearningZongsheng Cao, Anran Liu, Jun Xie, Lang Chen 等KDD 2026
