VLMimic: Vision Language Models are Visual Imitation Learner for Fine-grained Actions
Guangyan Chen, Meiling Wang, Te Cui, Yao Mu, Haoyang Lu, Tianxing Zhou, Zicai Peng, Mengxiao Hu, Haizhou Li, Li Yuan, Yi Yang, Yufeng Yue
Abstract
Visual imitation learning (VIL) provides an efficient and intuitive strategy for robotic systems to acquire novel skills. Recent advancements in Vision Language Models (VLMs) have demonstrated remarkable performance in vision and language reasoning capabilities for VIL tasks. Despite the progress, current VIL methods naively employ VLMs to learn high-level plans from human videos, relying on pre-defined motion primitives for executing physical interactions, which remains a major bottleneck. In this work, we present VLMimic, a novel paradigm that harnesses VLMs to directly learn even fine-grained action levels, only given a limited number of human videos. Specifically, VLMimic first grounds object-centric movements from human videos, and learns skills using hierarchical constraint representations, facilitating the derivation of skills with fine-grained action levels from limited human videos. These skills are refined and updated through an iterative comparison strategy, enabling efficient adaptation to unseen environments. Our extensive experiments exhibit that our VLMimic, using only 5 human videos, yields significant improvements of over 27% and 21% in RLBench and real-world manipulation tasks, and surpasses baselines by over 37% in long-horizon tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 669376af-08da-4c1f-94ed-4c1f4795b584Cited by top-tier papers4
- Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal LearningTianyi Bai, Yuxuan Fan, Jiantao Qiu, Fupeng Sun et al.NeurIPS 2025 · 12 citations
- GraphMimic: Graph-to-Graphs Generative Modeling from Videos for Policy LearningGuangyan Chen, Te Cui, Meiling Wang, Chengcai Yang et al.CVPR 2025
- RoboTwin: Dual-Arm Robot Benchmark with Generative Digital TwinsYao Mu, Tianxing Chen, Zanxin Chen, Shijia Peng et al.CVPR 2025
- Adaptive Data Augmentation with Multi-armed Bandit: Sample-Efficient Embedding Calibration for Implicit Pattern RecognitionMinxue Tang, Yangyang Yu, Aolin Ding, MAZIYAR BARAN POUYAN et al.CVPR 2026
Builds on12
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis et al.CVPR 2022 · 525 citations
- EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of ThoughtYao Mu, Qinglong Zhang, Mengkang Hu, Wenhai Wang et al.NeurIPS 2023 · 453 citations
- Associating Objects with Transformers for Video Object SegmentationZongxin Yang, Yunchao Wei, Yi YangNeurIPS 2021 · 398 citations
- Decoupling Features in Hierarchical Propagation for Video Object SegmentationZongxin Yang, Yi YangNeurIPS 2022 · 243 citations
- Going Denser with Open-Vocabulary Part SegmentationPeize Sun, Shoufa Chen, Chenchen Zhu, Fanyi Xiao et al.ICCV 2023 · 83 citations
Related papers
- AR-VRM: Imitating Human Motions for Visual Robot Manipulation with Analogical ReasoningDejie Yang, Zijing Zhao, Yang LiuICCV 2025
- Vision-Language Foundation Models as Effective Robot ImitatorsXinghang Li, Minghuan Liu, Hanbo Zhang, Cunjun Yu et al.ICLR 2024 · 375 citations
- Gentle Manipulation Policy Learning via Demonstrations from VLM Planned Atomic SkillsJiayu Zhou, Qiwei Wu, Jian Li, Zhe Chen et al.AAAI 2026 · 1 citation
- ManiLong-Shot: Interaction-Aware One-Shot Imitation Learning for Long-Horizon ManipulationZixuan Chen, Chongkai Gao, Lin Shao, Jieqi Shi et al.AAAI 2026 · 1 citation
- HAMSTER: Hierarchical Action Models for Open-World Robot ManipulationYi Li, Yuquan Deng, Jesse Zhang, Joel Jang et al.ICLR 2025 · 1 citation
