Masked Skill Token Training for Hierarchical Off-Dynamics Transfer
Zeyu Feng, Haiyan Yin, Yew-Soon Ong, Harold Soh
Abstract
Generalizing policies across environments with altered dynamics remains a key challenge in reinforcement learning, particularly in offline settings where direct interaction or fine-tuning is impractical. We introduce Masked Skill Token Training (MSTT), a fully offline hierarchical RL framework that enables policy transfer using observation-only demonstrations. MSTT constructs a discrete skill space via unsupervised trajectory tokenization and trains a skill-conditioned value function using masked Bellman updates, which simulate dynamics shifts by selectively disabling skills. A diffusion-based trajectory generator, paired with feasibility-based filtering, enables the agent to execute valid, temporally extended actions without requiring action labels or access to the target environment. Our results in both discrete and continuous domains demonstrate the potential of mask-guided planning for robust generalization under dynamics shifts. To our knowledge, MSTT is the first work to explore masking as a mechanism for simulating and generalizing across off-dynamics environments. It marks a promising step toward scalable, structure-aware transfer and opens avenues to explore multi-goal conditioning, and extensions to more complex, real-world scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on28
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra et al.ICCV 2019 · 1,863 citations
- Planning with Diffusion for Flexible Behavior SynthesisMichael Janner, Yilun Du, Joshua B. Tenenbaum, Sergey LevineICML 2022 · 1,115 citations
- Dynamics-Aware Unsupervised Discovery of SkillsArchit Sharma, Shixiang Gu, Sergey Levine, Vikash Kumar et al.ICLR 2020 · 475 citations
- Finite Scalar Quantization: VQ-VAE Made SimpleFabian Mentzer, David Minnen, Eirikur Agustsson, Michael TschannenICLR 2024 · 442 citations
- Habitat 3.0: A Co-Habitat for Humans, Avatars, and RobotsXavier Puig, Eric Undersander, Andrew Szot, Mikael Dallaire Cote et al.ICLR 2024 · 252 citations
Related papers
- Robust Policy Learning via Offline Skill DiffusionWoo Kyung Kim, Minjong Yoo, Honguk WooAAAI 2024 · 9 citations
- MetaDiffuser: Diffusion Model as Conditional Planner for Offline Meta-RLFei Ni, Jianye Hao, Yao Mu, Yifu Yuan et al.ICML 2023 · 75 citations
- Foundation Policies with Hilbert RepresentationsSeohong Park, Tobias Kreiman, Sergey LevineICML 2024 · 72 citations
- Masked Trajectory Models for Prediction, Representation, and ControlPhilipp Wu, Arjun Majumdar, Kevin Stone, Yixin Lin et al.ICML 2023 · 57 citations
- Goal-Oriented Skill Abstraction for Offline Multi-Task Reinforcement LearningJinmin He, Kai Li, Yifan Zang, Haobo Fu et al.ICML 2025
