VistaWise: Building Cost-Effective Agent with Cross-Modal Knowledge Graph for Minecraft
Honghao Fu, Junlong Ren, Qi Chai, Deheng Ye, Yujun Cai, Hao Wang
摘要
Large language models (LLMs) have shown significant promise in embodied decisionmaking tasks within virtual open-world environments. Nonetheless, their performance is hindered by the absence of domain-specific knowledge. Methods that finetune on largescale domain-specific data entail prohibitive development costs. This paper introduces Vista-Wise, a cost-effective agent framework that integrates cross-modal domain knowledge and finetunes a dedicated object detection model for visual analysis. It reduces the requirement for domain-specific training data from millions of samples to a few hundred. VistaWise integrates visual information and textual dependencies into a cross-modal knowledge graph (KG), enabling a comprehensive and accurate understanding of multimodal environments. We also equip the agent with a retrieval-based pooling strategy to extract task-related information from the KG, and a desktop-level skill library to support direct operation of the Minecraft desktop client via mouse and keyboard inputs. Experimental results demonstrate that VistaWise achieves state-of-the-art performance across various open-world tasks, highlighting its effectiveness in reducing development costs while enhancing agent performance. * The work was done during an internship at HKUST(GZ). † Corresponding Author. ‡ Equal Contribution. KG Construction Memory Stack Desktop-level Skill Library Skill 1 def mine(duration): pyautogui.mouseDown(button='left') time.sleep(duration) pyautogui.mouseUp(button='left')
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- ContextNav: Towards Agentic Multimodal In-Context LearningHonghao Fu, Yuan Ouyang, Kai-Wei Chang, Yiwei Wang 等ICLR 2026 · 被引用 14 次
- Test-Time Attention Purification for Backdoored Large Vision Language ModelsZhifang Zhang, Bojun Yang, Shuo He, Weitong Chen 等CVPR 2026 · 被引用 7 次
- Experience Transfer for Multimodal LLM Agents in Minecraft GameChenghao Li, Jun Liu, Songbo Zhang, Huadong Jian 等CVPR 2026 · 被引用 4 次
- VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAGHonghao Fu, Miao Xu, Yiwei Wang, Dailing Zhang 等ACL 2026 · 被引用 2 次
- ReCAPA: Hierarchical Predictive Correction to Mitigate Cascading FailuresXiyin Zeng, Yuyu Sun, Haoyang Li, Shouqiang Liu 等ICLR 2026 · 被引用 2 次
它引用的顶会 Paper14
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
- Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online VideosBowen Baker, Ilge Akkaya, Peter Zhokhov, Joost Huizinga 等NeurIPS 2022 · 被引用 458 次
- Active Retrieval Augmented GenerationZhengbao Jiang, Frank F. Xu, Luyu Gao, Zhiqing Sun 等EMNLP 2023 · 被引用 315 次
- Think-on-Graph: Deep and Responsible Reasoning of Large Language Model on Knowledge GraphJiashuo Sun, Chengjin Xu, Lumingyuan Tang, Saizhuo Wang 等ICLR 2024 · 被引用 247 次
- Subgraph Retrieval Enhanced Model for Multi-hop Knowledge Base Question AnsweringJing Zhang, Xiaokang Zhang, Jifan Yu, Jian Tang 等ACL 2022 · 被引用 221 次
相关 Paper
- GraphVis: Boosting LLMs with Visual Knowledge Graph IntegrationYihe Deng, Chenchen Ye, Zijie Huang, Mingyu Derek Ma 等NeurIPS 2024 · 被引用 23 次
- Connecting the Dots: Training-Free Visual Grounding via Agentic ReasoningLiqin Luo, Guangyao Chen, Xiawu Zheng, Yongxing Dai 等AAAI 2026
- Optimus-1: Hybrid Multimodal Memory Empowered Agents Excel in Long-Horizon TasksZaijing Li, Yuquan Xie, Rui Shao, Gongwei Chen 等NeurIPS 2024 · 被引用 104 次
- VELMA: Verbalization Embodiment of LLM Agents for Vision and Language Navigation in Street ViewRaphael Schumann, Wanrong Zhu, Weixi Feng, Tsu-Jui Fu 等AAAI 2024 · 被引用 122 次
- Embodied Multi-Modal Agent trained by an LLM from a Parallel TextWorldYijun Yang, Tianyi Zhou, Kanxue Li, Dapeng Tao 等CVPR 2024 · 被引用 23 次
