Skill Transformer: A Monolithic Policy for Mobile Manipulation
Xiaoyu Huang, Dhruv Batra, Akshara Rai, Andrew Szot
摘要
We present Skill Transformer, an approach for solving long-horizon robotic tasks by combining conditional sequence modeling and skill modularity. Conditioned on egocentric and proprioceptive observations of a robot, Skill Transformer is trained end-to-end to predict both a high-level skill (e.g., navigation, picking, placing), and a whole-body low-level action (e.g., base and arm motion), using a transformer architecture and demonstration trajectories that solve the full task. It retains the composability and modularity of the overall task through a skill predictor module while reasoning about low-level actions and avoiding hand-off errors, common in modular approaches. We test Skill Transformer on an embodied rearrangement benchmark and find it performs robust task planning and low-level control in new scenarios, achieving a 2.5x higher success rate than baselines in hard rearrangement problems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Large Language Models as Generalizable Policies for Embodied TasksAndrew Szot, Max Schwarzer, Harsh Agrawal, Bogdan Mazoure 等ICLR 2024 · 被引用 114 次
- ELMUR: External Layer Memory with Update/Rewrite for Long-Horizon RL ProblemsEgor Cherepanov, Alexey Kovalev, Aleksandr PanovICLR 2026 · 被引用 1 次
- MoManipVLA: Transferring Vision-language-action Models for General Mobile ManipulationZhenyu Wu, Yuheng Zhou, Xiuwei Xu, Ziwei Wang 等CVPR 2025
它引用的顶会 Paper16
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee 等NeurIPS 2021 · 被引用 2,557 次
- Habitat 2.0: Training Home Assistants to Rearrange their HabitatAndrew Szot, Alexander Clegg, Eric Undersander, Erik Wijmans 等NeurIPS 2021 · 被引用 826 次
- DD-PPO: Learning Near-Perfect PointGoal Navigators from 2.5 Billion FramesErik Wijmans, Abhishek Kadian, Ari Morcos, Stefan Lee 等ICLR 2020 · 被引用 608 次
- Dynamics-Aware Unsupervised Discovery of SkillsArchit Sharma, Shixiang Gu, Sergey Levine, Vikash Kumar 等ICLR 2020 · 被引用 475 次
- Multi-Game Decision TransformersKuang-Huei Lee, Ofir Nachum, Mengjiao Yang, Lisa Lee 等NeurIPS 2022 · 被引用 279 次
相关 Paper
- SkillDiffuser: Interpretable Hierarchical Planning via Skill Abstractions in Diffusion-Based Task ExecutionZhixuan Liang, Yao Mu, Hengbo Ma, Masayoshi Tomizuka 等CVPR 2024
- Structural Action Transformer for 3D Dexterous ManipulationXiaohan Lei, Min Wang, Bohong Weng, Wengang Zhou 等CVPR 2026
- Multi-skill Mobile Manipulation for Object RearrangementJiayuan Gu, Devendra Singh Chaplot, Hao Su, Jitendra MalikICLR 2023 · 被引用 10 次
- Chain-of-Thought Predictive ControlZhiwei Jia, Vineet Thumuluri, Fangchen Liu, Linghao Chen 等ICML 2024 · 被引用 24 次
- Breaking Down and Building Up: Mixture of Skill-Based Vision-and-Language Navigation AgentsTianyi Ma, Yue Zhang, Zehao Wang, Parisa KordjamshidiACL 2026 · 被引用 3 次
