Structural Action Transformer for 3D Dexterous Manipulation
Xiaohan Lei, Min Wang, Bohong Weng, Wengang Zhou, Houqiang Li
Abstract
Achieving human-level dexterity in robots via imitation learning from heterogeneous datasets is hindered by the challenge of cross-embodiment skill transfer, particularly for high-DoF robotic hands. Existing methods, often relying on 2D observations and temporal-centric action representation, struggle to capture 3D spatial relations and fail to handle embodiment heterogeneity. This paper proposes the Structural Action Transformer (SAT), a new 3D dexterous manipulation policy that challenges this paradigm by introducing a structural-centric perspective. We reframe each action chunk not as a temporal sequence, but as a variable-length, unordered sequence of joint-wise trajectories. This structural formulation allows a Transformer to natively handle heterogeneous embodiments, treating the joint count as a variable sequence length. To encode structural priors and resolve ambiguity, we introduce an Embodied Joint Codebook that embeds each joint's functional role and kinematic properties. Our model learns to generate these trajectories from 3D point clouds via a continuous-time flow matching objective. We validate our approach by pre-training on large-scale heterogeneous datasets and fine-tuning on simulation and real-world dexterous manipulation tasks. Our method consistently outperforms all baselines, demonstrating superior sample efficiency and effective cross-embodiment skill transfer. This structural-centric representation offers a new path toward scaling policies for high-DoF, heterogeneous manipulators.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ba57d059-c151-4b00-b567-d3a4c16b80e3Builds on28
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch et al.ICML 2023 · 2,601 citations
- Planning with Diffusion for Flexible Behavior SynthesisMichael Janner, Yilun Du, Joshua B. Tenenbaum, Sergey LevineICML 2022 · 1,115 citations
- Dynamics-Aware Unsupervised Discovery of SkillsArchit Sharma, Shixiang Gu, Sergey Levine, Vikash Kumar et al.ICLR 2020 · 475 citations
- 3D-VLA: A 3D Vision-Language-Action Generative World ModelHaoyu Zhen, Xiaowen Qiu, Peihao Chen, Jincheng Yang et al.ICML 2024 · 303 citations
Related papers
- Scaling Proprioceptive-Visual Learning with Heterogeneous Pre-trained TransformersLirui Wang, Xinlei Chen, Jialiang Zhao, Kaiming HeNeurIPS 2024 · 208 citations
- Skill Transformer: A Monolithic Policy for Mobile ManipulationXiaoyu Huang, Dhruv Batra, Akshara Rai, Andrew SzotICCV 2023 · 34 citations
- H-RDT: Human Manipulation Enhanced Bimanual Robotic ManipulationHongzhe Bi, Lingxuan Wu, Tianwei Lin, Hengkai Tan et al.AAAI 2026 · 25 citations
- X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action ModelJinliang Zheng, Jianxiong Li, Zhihao Wang, Dongxiu Liu et al.ICLR 2026 · 335 citations
- Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action PolicyZhi Hou, Tianyi Zhang, Yuwen Xiong, Haonan Duan et al.ICCV 2025 · 9 citations
