Hierarchical Procedural Meta-Reasoning for Generalizable Multimodal Agents
Yao Fu, Shengyi Qian, Pierluca D'Oro, Fanyi Xiao, Honglak Lee, Joseph Tighe, Manchen Wang
摘要
While multimodal agents can achieve strong performance through fine-tuning, their ability to generalize remains limited in complex real-world tasks such as mobile navigation, where diverse applications, frequent system changes, and customized workflows are common in practice. We argue that a fundamental bottleneck lies in whether an agent possesses sufficient task-specific procedural knowledge to accomplish a given goal. Such procedural knowledge may be provided by the general capabilities of large language models, or obtained from additional external resources such as web search when necessary. Based on this view, we propose Procedure-Aware Multimodal Agent with Meta Reasoning, a framework that explicitly represents task knowledge as natural-language procedures and trains a procedure-aware grounded agent to condition its actions on this knowledge. By learning to leverage procedural knowledge from different sources, our approach enables robust generalization across tasks, applications, interface versions, and multi-app workflows, achieving substantial improvements on challenging Android benchmarks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language ModelsChan Hee Song, Brian M. Sadler, Jiaman Wu, Wei-Lun Chao 等ICCV 2023 · 被引用 685 次
- GPT-4V(ision) is a Generalist Web Agent, if GroundedBoyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun 等ICML 2024 · 被引用 496 次
- What matters when building vision-language models?Hugo Laurençon, Léo Tronchon, Matthieu Cord, Victor SanhNeurIPS 2024 · 被引用 401 次
- Goal-Conditioned Reinforcement Learning with Imagined SubgoalsElliot Chane-Sane, Cordelia Schmid, Ivan LaptevICML 2021 · 被引用 183 次
- Multimodal Web Navigation with Instruction-Finetuned Foundation ModelsHiroki Furuta, Kuang-Huei Lee, Ofir Nachum, Yutaka Matsuo 等ICLR 2024 · 被引用 160 次
相关 Paper
- UINavBench: A Framework for Comprehensive Evaluation of Interactive Digital AgentsHarsh Agrawal, Eldon Schoop, Xinlei Pan, Anuj Mahajan 等ICCV 2025 · 被引用 9 次
- MAKAR: a Multi-Agent framework based Knowledge-Augmented Reasoning for Grounded Multimodal Named Entity RecognitionXinkui Lin, Yuhui Zhang, Yongxiu Xu, Kun Huang 等EMNLP 2025
- MAGNET: Towards Adaptive GUI Agents with Memory-Driven Knowledge EvolutionLibo Sun, Jiwen Zhang, Siyuan Wang, Zhongyu WeiACL 2026 · 被引用 5 次
- GUI-Rise: Structured Reasoning and History Summarization for GUI NavigationTao Liu, Chongyu Wang, Rongjie Li, Yingchen Yu 等NeurIPS 2025 · 被引用 4 次
- Agent-SAMA: State-Aware Mobile AssistantLinqiang Guo, Wei Liu, Yi Wen Heng, Tse-Hsun (Peter) Chen 等AAAI 2026 · 被引用 2 次
