3D Equivariant Visuomotor Policy Learning via Spherical Projection
Boce Hu, Dian Wang, David Klee, Heng Tian, Xupeng Zhu, Haojie Huang, Robert Platt, Robin Walters
Abstract
Equivariant models have recently been shown to improve the data efficiency of diffusion policy by a significant margin. However, prior work that explored this direction focused primarily on point cloud inputs generated by multiple cameras fixed in the workspace. This type of point cloud input is not compatible with the now-common setting where the primary input modality is an eye-in-hand RGB camera like a GoPro. This paper closes this gap by incorporating into the diffusion policy model a process that projects features from the 2D RGB camera image onto a sphere. This enables us to reason about symmetries in without explicitly reconstructing a point cloud. We perform extensive experiments in both simulation and the real world that demonstrate that our method consistently outperforms strong baselines in terms of both performance and sample efficiency. Our work, Image-to-Sphere Policy (), is the first -equivariant policy learning framework for robotic manipulation that works using only monocular RGB inputs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 62d1e7a2-689d-4349-8a16-e692c6b00aebCited by top-tier papers2
- Projective Equivariant Networks via Second-order Fundamental Differential InvariantsYikang Li, Yeqing Qiu, Yuxuan Chen, Lingshen He et al.NeurIPS 2025 · 1 citation
- RAVEN: End-to-end Equivariant Robot Learning with RGB CamerasDavid Klee, Boce Hu, Andrew Cole, Heng Tian et al.ICLR 2026
Builds on25
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
- Planning with Diffusion for Flexible Behavior SynthesisMichael Janner, Yilun Du, Joshua B. Tenenbaum, Sergey LevineICML 2022 · 1,115 citations
- Behavior Transformers: Cloning modes with one stoneNur Muhammad Shafiullah, Zichen Jeff Cui, Ariuntuya Altanzaya, Lerrel PintoNeurIPS 2022 · 470 citations
- A Program to Build E(N)-Equivariant Steerable CNNsGabriele Cesa, Leon Lang, Maurice WeilerICLR 2022 · 133 citations
Related papers
- SE(3)-Equivariant Diffusion Policy in Spherical Fourier SpaceXupeng Zhu, Fan Wang, Robin Walters, Jane ShiICML 2025
- Efficient Hybrid SE(3)-Equivariant Visuomotor Flow Policy via Spherical Harmonics for Robot ManipulationQinglun Zhang, Shen Cheng, Tian Dan, Haoqiang Fan et al.CVPR 2026 · 1 citation
- Image to Sphere: Learning Equivariant Features for Efficient Pose PredictionDavid Klee, Ondrej Biza, Robert Platt, Robin WaltersICLR 2023 · 3 citations
- A Practical Guide for Incorporating Symmetry in Diffusion PolicyDian Wang, Boce Hu, Shuran Song, Robin Walters et al.NeurIPS 2025 · 9 citations
- Equivariant Descriptor Fields: SE(3)-Equivariant Energy-Based Models for End-to-End Visual Robotic Manipulation LearningHyunwoo Ryu, Hong-in Lee, Jeong-Hoon Lee, Jongeun ChoiICLR 2023 · 11 citations
