Scaling up Memory for Robotic Control via Experience Retrieval
Ajay Sridhar, Jennifer Pan, Satvik Sharma, Chelsea Finn
摘要
Humans routinely rely on memory to perform tasks, yet most robot policies lack this capability; our goal is to endow robot policies with the same ability. Naively conditioning on long observation histories is computationally expensive and brittle under covariate shift, while indiscriminate subsampling of history leads to irrelevant or redundant information. We propose a hierarchical policy framework, where the high-level policy is trained to select and track previous relevant keyframes from its experience. The high-level policy uses selected keyframes and the most recent frames when generating text instructions for a lowlevel policy to execute. This design is compatible with existing vision-languageaction (VLA) models and enables the system to efficiently reason over longhorizon dependencies. In our experiments, we finetune Qwen2.5-VL-7B-Instruct and π 0.5 as the high-level and low-level policies respectively, using demonstrations supplemented with minimal language annotations. Our approach, MemER, outperforms prior methods on three real-world long-horizon robotic manipulation tasks that require minutes of memory. Videos and code can be found at https://jen-pan.github.io/memer/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- HiMe: Hierarchical Embodied Memory for Long-Horizon Vision-Language-Action ControlLi Ji, Siyin Wang, Pengfang Qian, Xiaopeng Yu 等ICML 2026 · 被引用 1 次
- CycleManip: Enabling Cycle-based Manipulation via Effective History Perception and UnderstandingYi-Lin Wei, Haoran Liao, Yuhao Lin, Pengyue Wang 等CVPR 2026
它引用的顶会 Paper8
- Self-Chained Image-Language Model for Video Localization and Question AnsweringShoubin Yu, Jaemin Cho, Prateek Yadav, Mohit BansalNeurIPS 2023 · 被引用 281 次
- Robust Fine-tuning of Vision-Language-Action Robot Policies via Parameter MergingYajat Yadav, Zhiyuan Zhou, Andrew Wagenmaker, Karl Pertsch 等ICLR 2026 · 被引用 10 次
- HAMSTER: Hierarchical Action Models for Open-World Robot ManipulationYi Li, Yuquan Deng, Jesse Zhang, Joel Jang 等ICLR 2025 · 被引用 1 次
- TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic PoliciesRuijie Zheng, Yongyuan Liang, Shuaiyi Huang, Jianfeng Gao 等ICLR 2025
- SAM2Act: Integrating Visual Foundation Model with A Memory Architecture for Robotic ManipulationHaoquan Fang, Markus Grotz, Wilbert Pumacay, Yi Ru Wang 等ICML 2025
相关 Paper
- MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic ManipulationHao Shi, Bin Xie, Yingfei Liu, Lin Sun 等ICLR 2026 · 被引用 227 次
- HAMLET: Switch Your Vision-Language-Action Model into a History-Aware PolicyMyungkyu Koo, Daewon Choi, Taeyoung Kim, Kyungmin Lee 等ICLR 2026 · 被引用 52 次
- TRM-VLA: Temporal-Aware Chain-of-Thought Reasoning and Memorization for Vision-Language-Action ModelsLI XIANG, Yali Li, Yuan Wang, Shengjin WangCVPR 2026
- SeqWalker: Sequential-Horizon Vision-and-Language Navigation with Hierarchical PlanningZebin Han, Xudong Wang, Baichen Liu, Qi Lyu 等AAAI 2026 · 被引用 2 次
- History Aware Multimodal Transformer for Vision-and-Language NavigationShizhe Chen, Pierre-Louis Guhur, Cordelia Schmid, Ivan LaptevNeurIPS 2021 · 被引用 427 次
