Retrieval-Augmented Embodied Agents
Yichen Zhu, Zhicai Ou, Xiaofeng Mou, Jian Tang
Abstract
Embodied agents operating in complex and uncertain environments face considerable challenges. While some advanced agents handle complex manipulation tasks with proficiency, their success often hinges on extensive training data to develop their capabilities. In contrast, humans typically rely on recalling past experiences and analogous situations to solve new problems. Aiming to emulate this human approach in robotics, we introduce the Retrieval-Augmented Embodied Agent (RAEA). This innovative system equips robots with a form of shared memory, significantly enhancing their performance. Our approach integrates a policy retriever, allowing robots to access relevant strategies from an external policy memory bank based on multi-modal inputs. Additionally, a policy generator is employed to assimilate these strategies into the learning process, enabling robots to formulate effective responses to tasks. Extensive testing of RAEA in both simulated and real-world scenarios demonstrates its superior performance over traditional methods, representing a major leap forward in robotic technology.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 156e6d07-ea80-48a3-bc36-e5300e990ca9Cited by top-tier papers11
- RAGraph: A General Retrieval-Augmented Graph Learning FrameworkXinke Jiang, Rihong Qiu, Yongxin Xu, Wentao Zhang et al.NeurIPS 2024 · 42 citations
- ChatVLA-2: Vision-Language-Action Model with Open-World ReasoningZhongyi Zhou, Yichen Zhu, Xiaoyu Liu, Zhibin Tang et al.NeurIPS 2025 · 10 citations
- How Do Multimodal Large Language Models Handle Complex Multimodal Reasoning? Placing Them in an Extensible Escape GameZiyue Wang, Yurui Dong, Fuwen Luo, Minyuan Ruan et al.ICCV 2025 · 9 citations
- From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action ModelBing Hu, Zaijing Li, Rui Shao, Junda Chen et al.ICML 2026 · 4 citations
- Dejavu: Towards Experience Feedback Learning for Embodied IntelligenceShaokai Wu, Yanbiao Ji, Qiuchang Li, Zhiyi Zhang et al.CVPR 2026 · 3 citations
Builds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- NExT-GPT: Any-to-Any Multimodal LLMShengqiong Wu, Hao Fei, Leigang Qu, Wei Ji et al.ICML 2024 · 786 citations
Related papers
- P-RAG: Progressive Retrieval Augmented Generation For Planning on Embodied Everyday TaskWeiye Xu, Min Wang, Wengang Zhou, Houqiang LiACM MM 2024 · 5 citations
- InstructRAG: Leveraging Retrieval-Augmented Generation on Instruction Graphs for LLM-Based Task PlanningZheng Wang, Shu Xian Teo, Jun Jie Chew, Wei ShiSIGIR 2025 · 4 citations
- MuRAG: Multimodal Retrieval-Augmented Generator for Open Question Answering over Images and TextWenhu Chen, Hexiang Hu, Xi Chen, Pat Verga et al.EMNLP 2022 · 89 citations
- RAPO: Expanding Exploration for LLM Agents via Retrieval-Augmented Policy OptimizationSiwei Zhang, Yun Xiong, Xi Chen, Zian Jia et al.KDD 2026 · 7 citations
- M-RAG: Reinforcing Large Language Model Performance through Retrieval-Augmented Generation with Multiple PartitionsZheng Wang, Shu Xian Teo, Jieer Ouyang, Yongjun Xu et al.ACL 2024
