Efficient Policy Adaptation with Contrastive Prompt Ensemble for Embodied Agents
Wonje Choi, Woo Kyung Kim, Seunghyun Kim, Honguk Woo
Abstract
For embodied reinforcement learning (RL) agents interacting with the environment, it is desirable to have rapid policy adaptation to unseen visual observations, but achieving zero-shot adaptation capability is considered as a challenging problem in the RL context. To address the problem, we present a novel contrastive prompt ensemble (ConPE) framework which utilizes a pretrained vision-language model and a set of visual prompts, thus enabling efficient policy learning and adaptation upon a wide range of environmental and physical changes encountered by embodied agents. Specifically, we devise a guided-attention-based ensemble approach with multiple visual prompts on the vision-language model to construct robust state representations. Each prompt is contrastively learned in terms of an individual domain factor that significantly affects the agent's egocentric perception and observation. For a given task, the attention-based ensemble and policy are jointly learned so that the resulting state representations not only generalize to various domains but are also optimized for learning the task. Through experiments, we show that ConPE outperforms other state-of-the-art algorithms for several embodied agent tasks including navigation in AI2THOR, manipulation in egocentric-Metaworld, and autonomous driving in CARLA, while also improving the sample efficiency of policy learning and adaptation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Embodied CoT Distillation From LLM To Off-the-shelf AgentsWonje Choi, Woo Kyung Kim, Minjong Yoo, Honguk WooICML 2024 · 13 citations
- LLM-Assisted Semantically Diverse Teammate Generation for Efficient Multi-agent CoordinationLihe Li, Lei Yuan, Pengsen Liu, Tao Jiang et al.ICML 2025
- Plan in Sandbox, Navigate in Open Worlds: Learning Physics-Grounded Abstracted Experience for Embodied NavigationZhixuan Shen, Jiawei Du, Ziyu Guo, Han Luo et al.ICML 2026
- MIRA: Memory-Integrated Reinforcement Learning Agent with Limited LLM GuidanceNarjes Nourzad, Carlee Joe-WongICLR 2026
- Cross-Domain Demo-to-Code via Neurosymbolic Counterfactual ReasoningJooyoung Kim, Wonje Choi, Younguk Song, Honguk WooCVPR 2026
Builds on17
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- CURL: Contrastive Unsupervised Representations for Reinforcement LearningMichael Laskin, Aravind Srinivas, Pieter AbbeelICML 2020 · 1,261 citations
- Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from PixelsDenis Yarats, Ilya Kostrikov, Rob FergusICLR 2021 · 911 citations
- Data-Efficient Reinforcement Learning with Self-Predictive RepresentationsMax Schwarzer, Ankesh Anand, Rishab Goel, R. Devon Hjelm et al.ICLR 2021 · 399 citations
- Decoupling Representation Learning from Reinforcement LearningAdam Stooke, Kimin Lee, Pieter Abbeel, Michael LaskinICML 2021 · 389 citations
Related papers
- ADAPT: Vision-Language Navigation with Modality-Aligned Action PromptsBingqian Lin, Yi Zhu, Zicong Chen, Xiwen Liang et al.CVPR 2022 · 45 citations
- Prompt-based Visual Alignment for Zero-shot Policy TransferHaihan Gao, Rui Zhang, Qi Yi, Hantao Yao et al.ICML 2024 · 1 citation
- Visual-Language Navigation Pretraining via Prompt-based Environmental Self-explorationXiwen Liang, Fengda Zhu, Lingling Li, Hang Xu et al.ACL 2022 · 37 citations
- PØDA: Prompt-driven Zero-shot Domain AdaptationMohammad Fahes, Tuan-Hung Vu, Andrei Bursuc, Patrick Pérez et al.ICCV 2023 · 82 citations
- Towards Difficulty-Agnostic Efficient Transfer Learning for Vision-Language ModelsYongjin Yang, Jongwoo Ko, Se-Young YunEMNLP 2024 · 1 citation
