The Lie We Tell: Correcting the Euclidean Fallacy in Vision Language Action Policies via Score Matching on Tangent Space
Bing-Cheng Chuang, I-Hsuan Chu, Bor Jiun Lin, Yang YuanFu, Min Sun, Chun-Yi Lee
摘要
Diffusion-based Vision-Language-Action policies achieve remarkable success in robotic manipulation, yet commit a fundamental geometric error we term the Euclidean Fallacy: representing SE(3) poses as flat vectors. This approximation induces (1) manifold drift violating SO(3) constraints, (2) broken equivariance under coordinate transformations, and (3) non-geodesic trajectories with excessive kinematic cost. We introduce Lie Diffuser Actor (LDA), a diffusion framework operating intrinsically on SE(3). Our method injects noise through left-invariant SDEs, predicts scores in the tangent space, and retracts samples via the exponential map. This formulation eliminates manifold drift by construction while guaranteeing coordinate-frame equivariance and geodesic optimality. On CALVIN ABCD, LDA improves average task length from to (). We further validate our method on real robot and the results show that our methodology outperforms the baseline on majority tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Riemannian Score-Based Generative ModellingValentin De Bortoli, Emile Mathieu, Michael J. Hutchinson, James Thornton 等NeurIPS 2022 · 被引用 306 次
- HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action ModelJiaming Liu, Hao Chen, Zhuoyang Liu, Pengju An 等ICLR 2026 · 被引用 216 次
- Flow Matching on General GeometriesRicky T. Q. Chen, Yaron LipmanICLR 2024 · 被引用 193 次
- Theseus: A Library for Differentiable Nonlinear OptimizationLuis Pineda, Taosha Fan, Maurizio Monge, Shobha Venkataraman 等NeurIPS 2022 · 被引用 124 次
相关 Paper
- SE(3)-Equivariant Diffusion Policy in Spherical Fourier SpaceXupeng Zhu, Fan Wang, Robin Walters, Jane ShiICML 2025
- Et-Seed: Efficient trajectory-Level SE(3) equivariant diffusion PolicyChenrui Tie, Yue Chen, Ruihai Wu, Boxuan Dong 等ICLR 2025
- A Primer on SO(3) Action Representations in Deep Reinforcement LearningMartin Schuck, Sherif Samy, Angela P. SchoelligICLR 2026 · 被引用 2 次
- EquAct: An SE(3)-Equivariant Multi-Task Transformer for 3D Robotic ManipulationXupeng Zhu, Yu Qi, Yizhe Zhu, Robin Walters 等ICLR 2026 · 被引用 9 次
- PDFactor: Learning Tri-Perspective View Policy Diffusion Field for Multi-Task Robotic ManipulationJingyi Tian, Le Wang, Sanping Zhou, Sen Wang 等CVPR 2025
