SAM-E: Leveraging Visual Foundation Model with Sequence Imitation for Embodied Manipulation
Junjie Zhang, Chenjia Bai, Haoran He, Zhigang Wang, Bin Zhao, Xiu Li, Xuelong Li
摘要
Acquiring a multi-task imitation policy in 3D manipulation poses challenges in terms of scene understanding and action prediction. Current methods employ both 3D representation and multi-view 2D representation to predict the poses of the robot's end-effector. However, they still require a considerable amount of high-quality robot trajectories, and suffer from limited generalization in unseen tasks and inefficient execution in long-horizon reasoning. In this paper, we propose SAM-E, a novel architecture for robot manipulation by leveraging a vision-foundation model for generalizable scene understanding and sequence imitation for long-term action reasoning. Specifically, we adopt Segment Anything (SAM) pre-trained on a huge number of images and promptable masks as the foundation model for extracting task-relevant features, and employ parameter-efficient fine-tuning on robot data for a better understanding of embodied scenarios. To address long-horizon reasoning, we develop a novel multi-channel heatmap that enables the prediction of the action sequence in a single pass, notably enhancing execution efficiency. Experimental results from various instruction-following tasks demonstrate that SAM-E achieves superior performance with higher execution efficiency compared to the baselines, and also significantly improves generalization in few-shot adaptation to new tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- LangScene-X: Reconstruct Generalizable 3D Language-Embedded Scenes with TriMap Video DiffusionFangfu Liu, Hao Li, Jiawei Chi, Hanyang Wang 等ICCV 2025 · 被引用 7 次
- SpatialActor: Exploring Disentangled Spatial Representations for Robust Robotic ManipulationHao Shi, Bin Xie, Yingfei Liu, Yang Yue 等AAAI 2026 · 被引用 2 次
- Focus-Then-Reuse: Fast Adaptation in Visual Perturbation EnvironmentsJiahui Wang, Chao Chen, Jiacheng Xu, Zongzhang Zhang 等NeurIPS 2025 · 被引用 1 次
- Training-Free Generation of Temporally Consistent Rewards from VLMsYinuo Zhao, Jiale Yuan, Zhiyuan Xu, Xiaoshuai Hao 等ICCV 2025 · 被引用 1 次
- Cortical Policy: A Dual-Stream View Transformer for Robotic ManipulationXuening Zhang, Qi Lv, Xiang Deng, Miao Zhang 等ICLR 2026 · 被引用 1 次
它引用的顶会 Paper34
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
相关 Paper
- MTSAM: Multi-Task Fine-Tuning for Segment Anything ModelXuehao Wang, Zhan Zhuang, Feiyang Ye, Yu ZhangICLR 2025
- Improving the Generalization of Segmentation Foundation Model under Distribution Shift via Weakly Supervised AdaptationHaojie Zhang, Yongyi Su, Xun Xu, Kui JiaCVPR 2024 · 被引用 26 次
- EfficientSAM: Leveraged Masked Image Pretraining for Efficient Segment AnythingYunyang Xiong, Bala Varadarajan, Lemeng Wu, Xiaoyu Xiang 等CVPR 2024 · 被引用 185 次
- InfoSAM: Fine-Tuning the Segment Anything Model from An Information-Theoretic PerspectiveYuanhong Zhang, Muyao Yuan, Weizhan Zhang, Tieliang Gong 等ICML 2025
- Parameter-Free Fine-tuning via Redundancy Elimination for Vision Foundation ModelsJiahuan Long, Tingsong Jiang, Wen Yao, Yizhe Xiong 等AAAI 2026
