Hierarchical Procedural Meta-Reasoning for Generalizable Multimodal Agents
Yao Fu, Shengyi Qian, Pierluca D'Oro, Fanyi Xiao, Honglak Lee, Joseph Tighe, Manchen Wang
Abstract
While multimodal agents can achieve strong performance through fine-tuning, their ability to generalize remains limited in complex real-world tasks such as mobile navigation, where diverse applications, frequent system changes, and customized workflows are common in practice. We argue that a fundamental bottleneck lies in whether an agent possesses sufficient task-specific procedural knowledge to accomplish a given goal. Such procedural knowledge may be provided by the general capabilities of large language models, or obtained from additional external resources such as web search when necessary. Based on this view, we propose Procedure-Aware Multimodal Agent with Meta Reasoning, a framework that explicitly represents task knowledge as natural-language procedures and trains a procedure-aware grounded agent to condition its actions on this knowledge. By learning to leverage procedural knowledge from different sources, our approach enables robust generalization across tasks, applications, interface versions, and multi-app workflows, achieving substantial improvements on challenging Android benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7df68b49-53ab-4509-83f9-01d6445c45bfBuilds on13
- LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language ModelsChan Hee Song, Brian M. Sadler, Jiaman Wu, Wei-Lun Chao et al.ICCV 2023 · 685 citations
- GPT-4V(ision) is a Generalist Web Agent, if GroundedBoyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun et al.ICML 2024 · 496 citations
- What matters when building vision-language models?Hugo Laurençon, Léo Tronchon, Matthieu Cord, Victor SanhNeurIPS 2024 · 401 citations
- Goal-Conditioned Reinforcement Learning with Imagined SubgoalsElliot Chane-Sane, Cordelia Schmid, Ivan LaptevICML 2021 · 183 citations
- Multimodal Web Navigation with Instruction-Finetuned Foundation ModelsHiroki Furuta, Kuang-Huei Lee, Ofir Nachum, Yutaka Matsuo et al.ICLR 2024 · 160 citations
Related papers
- UINavBench: A Framework for Comprehensive Evaluation of Interactive Digital AgentsHarsh Agrawal, Eldon Schoop, Xinlei Pan, Anuj Mahajan et al.ICCV 2025 · 9 citations
- MAKAR: a Multi-Agent framework based Knowledge-Augmented Reasoning for Grounded Multimodal Named Entity RecognitionXinkui Lin, Yuhui Zhang, Yongxiu Xu, Kun Huang et al.EMNLP 2025
- MAGNET: Towards Adaptive GUI Agents with Memory-Driven Knowledge EvolutionLibo Sun, Jiwen Zhang, Siyuan Wang, Zhongyu WeiACL 2026 · 5 citations
- GUI-Rise: Structured Reasoning and History Summarization for GUI NavigationTao Liu, Chongyu Wang, Rongjie Li, Yingchen Yu et al.NeurIPS 2025 · 4 citations
- Agent-SAMA: State-Aware Mobile AssistantLinqiang Guo, Wei Liu, Yi Wen Heng, Tse-Hsun (Peter) Chen et al.AAAI 2026 · 2 citations
