Latent Adaptation of Foundation Policies for Sim-to-Real Transfer
Longchao Da, Thirulogasankar Pranav Kutralingam, Lirong Xiang, Hua Wei
Abstract
The sim-to-real problem remains a critical challenge in the real-world application of reinforcement learning (RL). The conventional sim-to-real methods heavily rely on resource-intensive re-training of the policy network to adapt to new domains, which limits the flexibility of the deployment of RL policies in ever-changing environments. Inspired by human locomotion, where individuals adjust their gait to new surface conditions without relearning the skill of walking, we introduce Latent Adaptation of Foundation Policies (Found-adapt), a framework that decouples this problem into skill acquisition and environment adaptation. Our method first pretrains a foundation policy on unlabeled offline trajectories from the source simulator, capturing diverse long-horizon behaviors as reusable skills. At deployment, instead of retraining the policy, we perform efficient latent space adaptation: a small amount of target-domain data is collected to refine a latent representation through an adapter network that incorporates parameter efficient alignment, which produces a task-ready controller under various system dynamics. This adaptation occurs entirely in the latent space, avoiding costly policy optimization while enabling robust transfer. Empirical results across multiple locomotion tasks and dynamic variations demonstrate that our method significantly reduces the sim-to-real gap. Further sensitivity analysis provides interesting insights into the requirements for data quality and applicable situations. These findings highlight how foundation policies with latent adaptation could serve as a general and efficient paradigm for real-world RL deployment. Implementation and experiment are available here.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0c51e2b8-a694-420e-b9a3-d0e8bd6c726bBuilds on13
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Dynamics-Aware Unsupervised Discovery of SkillsArchit Sharma, Shixiang Gu, Sergey Levine, Vikash Kumar et al.ICLR 2020 · 475 citations
- Self-Supervised Policy Adaptation during DeploymentNicklas Hansen, Rishabh Jangir, Yu Sun, Guillem Alenyà et al.ICLR 2021 · 187 citations
- Discovering and Achieving Goals via World ModelsRussell Mendonca, Oleh Rybkin, Kostas Daniilidis, Danijar Hafner et al.NeurIPS 2021 · 177 citations
- Uncertainty-Aware Action Advising for Deep Reinforcement Learning AgentsFelipe Leno da Silva, Pablo Hernandez-Leal, Bilal Kartal, Matthew E. TaylorAAAI 2020 · 84 citations
Related papers
- Cross-modal Domain Adaptation for Cost-Efficient Visual Reinforcement LearningXiong-Hui Chen, Shengyi Jiang, Feng Xu, Zongzhang Zhang et al.NeurIPS 2021 · 15 citations
- Transferable Reinforcement Learning via Probabilistic Latent Embeddings and Dynamic Policy Adaptation for Sim-to-Real DeploymentGengyue Han, Yiheng FengICML 2026
- Domain Adaptation In Reinforcement Learning Via Latent Unified State RepresentationJinwei Xing, Takashi Nagata, Kexin Chen, Xinyun Zou et al.AAAI 2021 · 65 citations
- Towards Bridging the Gap between Large-Scale Pretraining and Efficient Finetuning for Humanoid ControlWeidong Huang, Zhehan Li, Hangxin Liu, Biao Hou et al.ICLR 2026 · 4 citations
- EgoBridge: Domain Adaptation for Generalizable Imitation from Egocentric Human DataRyan Punamiya, Dhruv Patel, Patcharapong Aphiwetsa, Pranav Kuppili et al.NeurIPS 2025 · 40 citations
