Move-Then-Operate: Behavioral Phasing for Human-Like Robotic Manipulation
Haoming Xu, Lei Lei, Jie Gu, Chu Tang, Jingmin Chen, Rui-Qi Wang
Abstract
We present Move-Then-Operate, a Vision language action framework that explicitly decouples robotic manipulation into two distinct behavioral phases: coarse relocation (move) and contactcritical interaction (operate). Unlike monolithic policies that conflate these heterogeneous regimes, our architecture employs a dual-expert policy routed by a learnable phase selector, introducing a structural inductive bias that isolates phasespecific dynamics. Phase labels are automatically generated via an MLLM-based pipeline conditioned on lightweight contextual cues such as endeffector velocity and subtask decomposition to ensure alignment with human motor patterns. Evaluated on the RoboTwin2 benchmark, our method achieves an average success rate of 68.9%, outperforming the monolithic π 0 baseline by 24%. It matches or exceeds models trained on 10× more data and reaches peak performance in 40% fewer training steps, demonstrating that architectural disentanglement of move and operate phases is a highly effective and efficient strategy for mastering high-precision manipulation. Work completed during the internship at Rightly Robotics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 57bd54f5-f712-4351-9659-e33a35c77811Builds on20
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch et al.ICML 2023 · 2,601 citations
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic ManipulationTianxing Chen, Zanxin Chen, Baijun Chen, Zijian Cai et al.ICML 2026 · 394 citations
- HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action ModelJiaming Liu, Hao Chen, Zhuoyang Liu, Pengju An et al.ICLR 2026 · 216 citations
- ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent PlanningChi-Pin Huang, Yueh-Hua Wu, Min-Hung Chen, Yu-Chiang Frank Wang et al.NeurIPS 2025 · 179 citations
Related papers
- Beyond Success: Refining Elegant Robot Manipulation from Mixed-Quality Data via Just-in-Time InterventionYanbo Mao, Jianlong Fu, Ruoxuan Zhang, Hongxia Xie et al.CVPR 2026 · 2 citations
- When would Vision-Proprioception Policies Fail in Robotic Manipulation?Jingxian Lu, Wenke Xia, Yuxuan Wu, Zhiwu Lu et al.ICLR 2026 · 11 citations
- Gentle Manipulation Policy Learning via Demonstrations from VLM Planned Atomic SkillsJiayu Zhou, Qiwei Wu, Jian Li, Zhe Chen et al.AAAI 2026 · 1 citation
- Rethinking Bimanual Robotic Manipulation: Learning with Decoupled Interaction FrameworkJian-Jian Jiang, Xiao-Ming Wu, Yi-Xiang He, Ling-An Zeng et al.ICCV 2025 · 4 citations
- MergeVLA: Cross-Skill Model Merging Toward a Generalist Vision-Language-Action AgentYuxia Fu, Zhizhen Zhang, Yuqi Zhang, Zijian Wang et al.CVPR 2026 · 21 citations
