AVIS: Autonomous Visual Information Seeking with Large Language Model Agent
Ziniu Hu, Ahmet Iscen, Chen Sun, Kai-Wei Chang, Yizhou Sun, David Ross, Cordelia Schmid, Alireza Fathi
摘要
In this paper, we propose an autonomous information seeking visual question answering framework, AVIS. Our method leverages a Large Language Model (LLM) to dynamically strategize the utilization of external tools and to investigate their outputs via tree search, thereby acquiring the indispensable knowledge needed to provide answers to the posed questions. Responding to visual questions that necessitate external knowledge, such as "What event is commemorated by the building depicted in this image?", is a complex task. This task presents a combinatorial search space that demands a sequence of actions, including invoking APIs, analyzing their responses, and making informed decisions. We conduct a user study to collect a variety of instances of human decision-making when faced with this task. This data is then used to design a system comprised of three components: an LLM-powered planner that dynamically determines which tool to use next, an LLM-powered reasoner that analyzes and extracts key information from the tool outputs, and a working memory component that retains the acquired information throughout the process. The collected user behavior serves as a guide for our system in two key ways. First, we create a transition graph by analyzing the sequence of decisions made by users. This graph delineates distinct states and confines the set of actions available at each state. Second, we use examples of user decision-making to provide our LLM-powered planner and reasoner with relevant contextual instances, enhancing their capacity to make informed decisions. We show that AVIS achieves state-of-the-art results on knowledge-intensive visual question answering benchmarks such as Infoseek [7] and OK-VQA [26]. * This work was done when Ziniu was an intern at Google. 37th Conference on Neural Information Processing Systems (NeurIPS 2023).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- SceneCraft: An LLM Agent for Synthesizing 3D Scenes as Blender CodeZiniu Hu, Ahmet Iscen, Aashi Jain, Thomas Kipf 等ICML 2024 · 被引用 105 次
- MMSearch-R1: Incentivizing LMMs to SearchJinming Wu, Zihao Deng, Wei Li, Yiding Liu 等ACL 2026 · 被引用 93 次
- Uncertainty of Thoughts: Uncertainty-Aware Planning Enhances Information Seeking in LLMsZhiyuan Hu, Chumin Liu, Xidong Feng, Yilun Zhao 等NeurIPS 2024 · 被引用 43 次
- MoReVQA: Exploring Modular Reasoning Models for Video Question AnsweringJuhong Min, Shyamal Buch, Arsha Nagrani, Minsu Cho 等CVPR 2024 · 被引用 27 次
- ToolVQA: A Dataset for Multi-Step Reasoning VQA with External ToolsShaofeng Yin, Ting Lei, Yang LiuICCV 2025 · 被引用 2 次
它引用的顶会 Paper24
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
相关 Paper
- TOA: Task-oriented Active VQAXiaoying Xing, Mingfu Liang, Ying WuNeurIPS 2023 · 被引用 20 次
- VizGenie: Toward Self-Refining, Domain-Aware Workflows for Next-Generation Scientific VisualizationAyan Biswas, Terece L. Turton, Nishath Rajiv Ranasinghe, Shawn M. Jones 等IEEE VIS 2025 · 被引用 1 次
- AvaTaR: Optimizing LLM Agents for Tool Usage via Contrastive ReasoningShirley Wu, Shiyu Zhao, Qian Huang, Kexin Huang 等NeurIPS 2024 · 被引用 95 次
- Lightweight Adaptive Topological Layout and Semantic Mapping in Vision-and-Language Navigation on WebsitesPingrui Lai, Zihao Xie, Hua YangAAAI 2026
- VideoPro: Adaptive Program Reasoning for Long Video UnderstandingChenglin Li, Feng Han, Yikun Wang, Ruilin Li 等ACL 2026 · 被引用 4 次
