Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation
Zaijing Li, Bing Hu, Rui Shao, Gongwei Chen, Dongmei Jiang, Pengwei Xie, Jianye Hao, Liqiang Nie
Abstract
Hierarchical Vision-Language-Action (VLA) models have rapidly become a dominant paradigm for robotic manipulation. It typically comprising a Vision-Language backbone for perception and understanding, together with a generative policy for action generation. However, its performance is increasingly bottlenecked by the action generation proceess. (i) Low inference efficiency. A pronounced distributional gap between isotropic noise priors and target action distributions, which increases denoising steps and the incidence of infeasible samples. (ii) Poor robustness. Existing policies condition solely on the current observation, neglecting the constraint of history sequence and thus lacking awareness of task progress and temporal consistency. To address these issues, we introduce OptimusVLA, a dual-memory VLA framework with Global Prior Memory (GPM) and Local Consistency Memory (LCM). GPM replaces Gaussian noise with task-level priors retrieved from semantically similar trajectories, thereby shortening the generative path and reducing the umber of function evaluations (NFE). LCM dynamically models executed action sequence to infer task progress and injects a learned consistency constraint that enforces temporal coherence and smoothness of trajectory. Across three simulation benchmarks, OptimusVLA consistently outperforms strong baselines: it achieves 98.6% average success rate on LIBERO, improves over pi_0 by 13.5% on CALVIN, and attains 38% average success rate on RoboTwin 2.0 Hard. In Real-World evaluation, OptimusVLA ranks best on Generalization and Long-horizon suites, surpassing pi_0 by 42.9% and 52.4%, respectively, while delivering 2.9x inference speedup.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ad2409de-6409-4ea8-bfb6-87a44e8134b0Cited by top-tier papers5
- ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic ManipulationWei Li, Jizhihui Liu, Yixing Li, Junwen Tong et al.CVPR 2026 · 8 citations
- HATS: Hardness-Aware Trajectory Synthesis for GUI AgentsRui Shao, Ruize Gao, Bin Xie, Yixing Li et al.CVPR 2026 · 7 citations
- From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action ModelBing Hu, Zaijing Li, Rui Shao, Junda Chen et al.ICML 2026 · 4 citations
- TINA: Text-Free Inversion Attack for Unlearned Text-to-Image Diffusion ModelsQianlong Xiang, Miao Zhang, Haoyu Zhang, Kun Wang et al.CVPR 2026 · 2 citations
- EnergyAction: Unimanual to Bimanual Composition with Energy-Based ModelsMingchen Song, Xiang Deng, Jie Wei, Dongmei Jiang et al.CVPR 2026 · 1 citation
Builds on32
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic ManipulationTianxing Chen, Zanxin Chen, Baijun Chen, Zijian Cai et al.ICML 2026 · 394 citations
- Vision-Language Foundation Models as Effective Robot ImitatorsXinghang Li, Minghuan Liu, Hanbo Zhang, Cunjun Yu et al.ICLR 2024 · 375 citations
- An Embodied Generalist Agent in 3D WorldJiangyong Huang, Silong Yong, Xiaojian Ma, Xiongkun Linghu et al.ICML 2024 · 361 citations
- DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World KnowledgeWenyao Zhang, Hongsi Liu, Zekun Qi, Yunnan Wang et al.NeurIPS 2025 · 244 citations
Related papers
- TRM-VLA: Temporal-Aware Chain-of-Thought Reasoning and Memorization for Vision-Language-Action ModelsLI XIANG, Yali Li, Yuan Wang, Shengjin WangCVPR 2026
- Unifying Perception and Action: A Hybrid-Modality Pipeline with Implicit Visual Chain-of-Thought for Robotic Action GenerationXiangkai Ma, Lekai Xing, Han Zhang, Wenzhong Li et al.CVPR 2026 · 11 citations
- HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action ModelsMinghui Lin, Pengxiang Ding, Shu Wang, Zifeng Zhuang et al.CVPR 2026 · 35 citations
- Evo-1: Lightweight Vision-Language-Action Model with Preserved Semantic AlignmentTao Lin, Yilei Zhong, Yuxin Du, Jingjing Zhang et al.CVPR 2026 · 47 citations
- On Robustness of Vision-Language-Action Model against Multi-Modal PerturbationsJianing Guo, Zhenhong Wu, Chang Tu, Yiyao Ma et al.ICLR 2026 · 7 citations
