EquAct: An SE(3)-Equivariant Multi-Task Transformer for 3D Robotic Manipulation
Xupeng Zhu, Yu Qi, Yizhe Zhu, Robin Walters, Robert Platt
摘要
Transformer architectures can effectively learn language-conditioned, multi-task 3D open-loop manipulation policies from demonstrations by jointly processing natural language instructions and 3D observations. However, although both the robot policy and language instructions inherently encode rich 3D geometric structures, standard transformers lack built-in guarantees of geometric consistency, often resulting in unpredictable behavior under SE(3) transformations of the scene. In this paper, we leverage SE(3) equivariance as a key structural property shared by both policy and language, and propose EquAct-a novel SE(3)-equivariant multi-task transformer. EquAct is theoretically guaranteed to be SE(3) equivariant and consists of two key components: (1) an efficient SE(3)-equivariant point cloud-based U-net with spherical Fourier features for policy reasoning, and (2) SE(3)-invariant Featurewise Linear Modulation (iFiLM) layers for language conditioning. To evaluate its spatial generalization ability, we benchmark EquAct on 18 RLBench simulation tasks with both SE(3) and SE(2) scene perturbations, and on 4 physical tasks. EquAct performs state-of-the-art across these simulation and physical tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- SE(3)-Transformers: 3D Roto-Translation Equivariant Attention NetworksFabian Fuchs, Daniel E. Worrall, Volker Fischer, Max WellingNeurIPS 2020 · 被引用 1,025 次
- Reinforcement Learning with Augmented DataMichael Laskin, Kimin Lee, Adam Stooke, Lerrel Pinto 等NeurIPS 2020 · 被引用 833 次
- Vector Neurons: A General Framework for SO(3)-Equivariant NetworksCongyue Deng, Or Litany, Yueqi Duan, Adrien Poulenard 等ICCV 2021 · 被引用 411 次
相关 Paper
- SE(3)-Equivariant Diffusion Policy in Spherical Fourier SpaceXupeng Zhu, Fan Wang, Robin Walters, Jane ShiICML 2025
- RAVEN: End-to-end Equivariant Robot Learning with RGB CamerasDavid Klee, Boce Hu, Andrew Cole, Heng Tian 等ICLR 2026
- Geometry-aware RL for Manipulation of Varying Shapes and Deformable ObjectsTai Hoang, Huy Le, Philipp Becker, Ngo Anh Vien 等ICLR 2025
- Deep SE(3)-Equivariant Geometric Reasoning for Precise Placement TasksBen Eisner, Yi Yang, Todor Davchev, Mel Vecerík 等ICLR 2024 · 被引用 24 次
- SE(3) Equivariant Convolution and Transformer in Ray SpaceYinshuang Xu, Jiahui Lei, Kostas DaniilidisNeurIPS 2023 · 被引用 6 次
