EnergyAction: Unimanual to Bimanual Composition with Energy-Based Models
Mingchen Song, Xiang Deng, Jie Wei, Dongmei Jiang, Liqiang Nie, Weili Guan
摘要
Recent advances in unimanual manipulation policies have achieved remarkable success across diverse robotic tasks through abundant training data and well-established model architectures. However, extending these capabilities to bimanual manipulation remains challenging due to the lack of bimanual demonstration data and the complexity of coordinating dual-arm actions. Existing approaches either rely on extensive bimanual datasets or fail to effectively leverage pre-trained unimanual policies. To address this limitation, we propose EnergyAction, a novel framework that compositionally transfers unimanual manipulation policies to bimanual tasks through the Energy-Based Models (EBMs). Specifically, our method incorporates three key innovations. First, we model individual unimanual policies as EBMs and leverage their compositional properties to compose left and right arm actions, enabling the fusion of unimanual policies into a bimanual policy. Second, we introduce an energy-based temporal-spatial coordination mechanism through energy constraints, ensuring the generated bimanual actions are both temporal coherence and spatial feasibility. Third, we propose two different energy-aware denoising strategies that dynamically adapt denoising steps based on action quality assessment. These strategies ensure the generation of high-quality actions while maintaining superior computational efficiency compared to fixed-step denoising approaches. Experimental results demonstrate that EnergyAction effectively transfers unimanual knowledge to bimanual tasks, achieving superior performance on both simulated and real-world tasks with minimal bimanual data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic ManipulationZaijing Li, Bing Hu, Rui Shao, Gongwei Chen 等CVPR 2026 · 被引用 23 次
- EnsembleVLA: Ensemble Learning for Vision-Language Action ModelsMingchen Song, Xiang Deng, Jie Wei, Dongmei Jiang 等ICML 2026
它引用的顶会 Paper35
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 被引用 11,724 次
- Compositional Visual Generation with Energy Based ModelsYilun Du, Shuang Li, Igor MordatchNeurIPS 2020 · 被引用 225 次
- QueST: Self-Supervised Skill Abstractions for Learning Continuous ControlAtharva Mete, Haotian Xue, Albert Wilcox, Yongxin Chen 等NeurIPS 2024 · 被引用 76 次
- Multi-Factor Adaptive Vision Selection for Egocentric Video Question AnsweringHaoyu Zhang, Meng Liu, Zixin Liu, Xuemeng Song 等ICML 2024 · 被引用 23 次
相关 Paper
- AnyBimanual: Transferring Unimanual Policy for General Bimanual ManipulationGuanxing Lu, Tengbo Yu, Haoyuan Deng, Season Si Chen 等ICCV 2025 · 被引用 2 次
- TwinVLA: Data-Efficient Bimanual Manipulation with Twin Single-Arm Vision-Language-Action ModelsHokyun Im, Euijin Jeong, Andrey Kolobov, Jianlong Fu 等ICLR 2026 · 被引用 11 次
- RDT-1B: a Diffusion Foundation Model for Bimanual ManipulationSongming Liu, Lingxuan Wu, Bangguo Li, Hengkai Tan 等ICLR 2025
- EMFuse: Energy-based Model Fusion for Decision MakingKejie He, Yi-Chen Li, Yang YuICLR 2026
- Action-Geometry Prediction with 3D Geometric Prior for Bimanual ManipulationChongyang Xu, Haipeng Li, Shen Cheng, Haoqiang Fan 等CVPR 2026 · 被引用 10 次
