VLM Agents Generate Their Own Memories: Distilling Experience into Embodied Programs of Thought
Gabriel Sarch, Lawrence Jang, Michael J. Tarr, William W. Cohen, Kenneth Marino, Katerina Fragkiadaki
摘要
Large-scale generative language and vision-language models (LLMs and VLMs) excel in few-shot learning but require high-quality demonstrations. We propose In-Context Abstraction Learning (ICAL), enabling VLM agents to transform suboptimal trajectories into high-quality training data through self-reflection and human feedback. Given imperfect task demonstrations, a VLM abstracts trajectories into generalized strategies and action annotations by correcting inefficiencies and annotating cognitive abstractions: causal relationships, object state changes, temporal subgoals, and task-relevant visual elements. These annotations are iteratively refined through human feedback during execution in similar environments. The resulting examples significantly improve decision-making when used for retrieval-augmented generation or fine-tuning. As the agent's example library grows, it becomes more efficient at abstracting new examples, requiring less human feedback and fewer environment interactions. ICAL achieves state-of-the-art results across multiple benchmarks. In TEACh dialogue-based instruction following, combining fine-tuning and retrieval on ICAL examples outperforms raw human demonstrations and expert examples by 17.5% in goal-condition success. In VisualWebArena, retrieval-augmented GPT-4V with ICAL improves task success 1.6x, while fine-tuned Qwen2-VL achieves 2.8x improvement over the base model. In Ego4D action forecasting, we surpass few-shot GPT-4V and remain competitive with supervised models. Our approach scales 2x better than raw demonstrations and significantly reduces manual prompt engineering requirements.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Grounded Reinforcement Learning for Visual ReasoningGabriel Sarch, Snigdha Saha, Naitik Khandelwal, Ayush Jain 等NeurIPS 2025 · 被引用 90 次
- GUI-explorer: Autonomous Exploration and Mining of Transition-aware Knowledge for GUI AgentBin Xie, Rui Shao, Gongwei Chen, Kaiwen Zhou 等ACL 2025 · 被引用 28 次
- IAG: Input-aware Backdoor Attack on VLM-based Visual GroundingJunxian Li, Beining Xu, Simin Chen, Jiatong Li 等CVPR 2026 · 被引用 13 次
- WALT: Web Agents that Learn ToolsViraj Prabhu, Yutong Dai, Matthew Fernandez, Krithika Ramakrishnan 等ICLR 2026 · 被引用 13 次
- Let's Think in Two Steps: Mitigating Agreement Bias in MLLMs with Self-Grounded VerificationMoises Andrade, Joonhyuk Cha, Brandon Ho, Vriksha Srihari 等ICLR 2026 · 被引用 12 次
它引用的顶会 Paper31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
- WebArena: A Realistic Web Environment for Building Autonomous AgentsShuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou 等ICLR 2024 · 被引用 1,197 次
相关 Paper
- Provoking Multi-modal Few-Shot LVLM via Exploration-Exploitation In-Context LearningCheng Chen, Yunpeng Zhai, Yifan Zhao, Jinyang Gao 等CVPR 2025
- VL-ICL Bench: The Devil in the Details of Multimodal In-Context LearningYongshuo Zong, Ondrej Bohdal, Timothy M. HospedalesICLR 2025
- Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement LearningSimon Zhai, Hao Bai, Zipeng Lin, Jiayi Pan 等NeurIPS 2024 · 被引用 214 次
- ExACT: Teaching AI Agents to Explore with Reflective-MCTS and Exploratory LearningXiao Yu, Baolin Peng, Vineeth Vajipey, Hao Cheng 等ICLR 2025
- NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language ModelsGengze Zhou, Yicong Hong, Qi WuAAAI 2024 · 被引用 361 次
