ManipLLM: Embodied Multimodal Large Language Model for Object-Centric Robotic Manipulation
Xiaoqi Li, Mingxu Zhang, Yiran Geng, Haoran Geng, Yuxing Long, Yan Shen, Renrui Zhang, Jiaming Liu, Hao Dong
摘要
Robot manipulation relies on accurately predicting contact points and end-effector directions to ensure successful operation. However, learning-based robot manipulation, trained on a limited category within a simulator, often struggles to achieve generalizability, especially when confronted with extensive categories. Therefore, we introduce an innovative approach for robot manipulation that leverages the robust reasoning capabilities of Multimodal Large Language Models (MLLMs) to enhance the stability and generalization of manipulation. By fine-tuning the injected adapters, we preserve the inherent common sense and reasoning ability of the MLLMs while equipping them with the ability for manipulation. The fundamental insight lies in the introduced fine-tuning paradigm, encompassing object category understanding, affordance prior reasoning, and object-centric pose prediction to stimulate the reasoning ability of MLLM in manipulation. During inference, our approach utilizes an RGB image and text prompt to predict the end effector's pose in chain of thoughts. After the initial contact is established, an active impedance adaptation policy is introduced to plan the upcoming waypoints in a closed-loop manner. Moreover, in real world, we design a test-time adaptation (TTA) strategy for manipulation to enable the model better adapt to the current real-world scene configuration. Experiments in simulator and real-world show the promising performance of Mani-pLLM. More details and demonstrations can be found at https://sites.google.com/view/manipllm.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper55
- DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World KnowledgeWenyao Zhang, Hongsi Liu, Zekun Qi, Yunnan Wang 等NeurIPS 2025 · 被引用 244 次
- HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action ModelJiaming Liu, Hao Chen, Zhuoyang Liu, Pengju An 等ICLR 2026 · 被引用 216 次
- Fast-in-Slow: A Dual-System VLA Model Unifying Fast Manipulation within Slow ReasoningHao Chen, Jiaming Liu, Chenyang Gu, Zhuoyang Liu 等NeurIPS 2025 · 被引用 74 次
- SoFar: Language-Grounded Orientation Bridges Spatial Reasoning and Object ManipulationZekun Qi, Wenyao Zhang, Yufei Ding, Runpei Dong 等NeurIPS 2025 · 被引用 65 次
- OmniNav: A Unified Framework for Prospective Exploration and Visual-Language NavigationXinda Xue, Junjun Hu, Minghua Luo, Xie Shichao 等ICLR 2026 · 被引用 51 次
它引用的顶会 Paper19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
相关 Paper
- RoboMP2: A Robotic Multimodal Perception-Planning Framework with Multimodal Large Language ModelsQi Lv, Hao Li, Xiang Deng, Rui Shao 等ICML 2024 · 被引用 4 次
- 3D-AffordanceLLM: Harnessing Large Language Models for Open-Vocabulary Affordance Detection in 3D WorldsHengshuo Chu, Xiang Deng, Qi Lv, Xiaoyang Chen 等ICLR 2025
- AffordBot: 3D Fine-grained Embodied Reasoning via Multimodal Large Language ModelsXinyi Wang, Xun Yang, Yanlong Xu, Yuchen Wu 等NeurIPS 2025 · 被引用 18 次
- Instruction-Augmented Long-Horizon Planning: Embedding Grounding Mechanisms in Embodied Mobile ManipulationFangyuan Wang, Shipeng Lyu, Peng Zhou, Anqing Duan 等AAAI 2025 · 被引用 9 次
- Zero-Shot Robotic Manipulation via 3D Gaussian Splatting-Enhanced Multimodal Retrieval-Augmented GenerationZilong Xie, Jingyu Gong, Xin Tan, Zhizhong Zhang 等AAAI 2026
