Failure is Feedback: History-Aware Backtracking for Agentic Traversal in Multimodal Graphs
Joohyung Yun, Doyup Lee, Wook-Shin Han
Abstract
Open-domain multimodal document retrieval aims to retrieve specific components (paragraphs, tables, or images) from large and interconnected document corpora. Existing graphbased retrieval approaches typically rely on a uniform similarity metric that overlooks hopspecific semantics, and their rigid pre-defined plans hinder dynamic error correction. These limitations suggest that a retriever should adapt its reasoning to the evolving context and recover intelligently from dead ends. To address these needs, we propose FAILURE IS FEED-BACK (FIF), which casts subgraph retrieval as a sequential decision process and introduces two key innovations. (i) We introduce a historyaware backtracking mechanism; unlike standard backtracking that simply reverts the state, our approach piggybacks on the context of failed traversals, leveraging insights from previous failures. (ii) We implement an economicallyrational agentic workflow. Unlike conventional agents with static strategies, our orchestrator employs a cost-aware traversal method to dynamically manage the trade-off between retrieval accuracy and inference costs, escalating to intensive LLM-based reasoning only when the prior failure justifies the additional computational investment. Extensive experiments show that FIF achieves state-of-the-art retrieval on the benchmarks of MULTIMODALQA, MMCOQA and
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6262707a-e1c9-41ac-a64f-7eb0c9f779b2Builds on12
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil et al.ICLR 2024 · 1,798 citations
- Reasoning on Graphs: Faithful and Interpretable Large Language Model ReasoningLinhao Luo, Yuan-Fang Li, Gholamreza Haffari, Shirui PanICLR 2024 · 499 citations
- Learning to Retrieve Reasoning Paths over Wikipedia Graph for Question AnsweringAkari Asai, Kazuma Hashimoto, Hannaneh Hajishirzi, Richard Socher et al.ICLR 2020 · 322 citations
- Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step QuestionsHarsh Trivedi, Niranjan Balasubramanian, Tushar Khot, Ashish SabharwalACL 2023 · 187 citations
- Hierarchical Graph Network for Multi-hop Question AnsweringYuwei Fang, Siqi Sun, Zhe Gan, Rohit Pillai et al.EMNLP 2020 · 157 citations
Related papers
- LILaC: Late Interacting in Layered Component Graph for Open-domain Multimodal Multihop RetrievalJoohyung Yun, Doyup Lee, Wook-Shin HanEMNLP 2025
- EvoGraph-R1: Self-Evolving Multimodal Knowledge Hypergraphs for Agentic RetrievalJiashi Lin, Changhong Jiang, Xiangru Lin, Ruifei Zhang et al.CVPR 2026 · 2 citations
- GraphRAG-R1: Graph Retrieval-Augmented Generation with Process-Constrained Reinforcement LearningChuanyue Yu, Kuo Zhao, Yuhan Li, Heng Chang et al.WWW 2026 · 8 citations
- HiKEY: Hierarchical Multimodal Retrieval for Open-Domain Document Question AnsweringJoongmin Shin, Gyuho Shim, Jeongbae Park, Jaehyung Seo et al.ACL 2026
- Towards Open-World Retrieval-Augmented Generation on Knowledge Graph: A Multi-Agent Collaboration FrameworkJiasheng Xu, Mingda Li, Yongqiang Tang, Peijie Wang et al.WWW 2026
