Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation
Zaijing Li, Bing Hu, Rui Shao, Gongwei Chen, Dongmei Jiang, Pengwei Xie, Jianye Hao, Liqiang Nie
摘要
Hierarchical Vision-Language-Action (VLA) models have rapidly become a dominant paradigm for robotic manipulation. It typically comprising a Vision-Language backbone for perception and understanding, together with a generative policy for action generation. However, its performance is increasingly bottlenecked by the action generation proceess. (i) Low inference efficiency. A pronounced distributional gap between isotropic noise priors and target action distributions, which increases denoising steps and the incidence of infeasible samples. (ii) Poor robustness. Existing policies condition solely on the current observation, neglecting the constraint of history sequence and thus lacking awareness of task progress and temporal consistency. To address these issues, we introduce OptimusVLA, a dual-memory VLA framework with Global Prior Memory (GPM) and Local Consistency Memory (LCM). GPM replaces Gaussian noise with task-level priors retrieved from semantically similar trajectories, thereby shortening the generative path and reducing the umber of function evaluations (NFE). LCM dynamically models executed action sequence to infer task progress and injects a learned consistency constraint that enforces temporal coherence and smoothness of trajectory. Across three simulation benchmarks, OptimusVLA consistently outperforms strong baselines: it achieves 98.6% average success rate on LIBERO, improves over pi_0 by 13.5% on CALVIN, and attains 38% average success rate on RoboTwin 2.0 Hard. In Real-World evaluation, OptimusVLA ranks best on Generalization and Long-horizon suites, surpassing pi_0 by 42.9% and 52.4%, respectively, while delivering 2.9x inference speedup.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic ManipulationWei Li, Jizhihui Liu, Yixing Li, Junwen Tong 等CVPR 2026 · 被引用 8 次
- HATS: Hardness-Aware Trajectory Synthesis for GUI AgentsRui Shao, Ruize Gao, Bin Xie, Yixing Li 等CVPR 2026 · 被引用 7 次
- From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action ModelBing Hu, Zaijing Li, Rui Shao, Junda Chen 等ICML 2026 · 被引用 4 次
- TINA: Text-Free Inversion Attack for Unlearned Text-to-Image Diffusion ModelsQianlong Xiang, Miao Zhang, Haoyu Zhang, Kun Wang 等CVPR 2026 · 被引用 2 次
- EnergyAction: Unimanual to Bimanual Composition with Energy-Based ModelsMingchen Song, Xiang Deng, Jie Wei, Dongmei Jiang 等CVPR 2026 · 被引用 1 次
它引用的顶会 Paper32
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic ManipulationTianxing Chen, Zanxin Chen, Baijun Chen, Zijian Cai 等ICML 2026 · 被引用 394 次
- Vision-Language Foundation Models as Effective Robot ImitatorsXinghang Li, Minghuan Liu, Hanbo Zhang, Cunjun Yu 等ICLR 2024 · 被引用 375 次
- An Embodied Generalist Agent in 3D WorldJiangyong Huang, Silong Yong, Xiaojian Ma, Xiongkun Linghu 等ICML 2024 · 被引用 361 次
- DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World KnowledgeWenyao Zhang, Hongsi Liu, Zekun Qi, Yunnan Wang 等NeurIPS 2025 · 被引用 244 次
相关 Paper
- TRM-VLA: Temporal-Aware Chain-of-Thought Reasoning and Memorization for Vision-Language-Action ModelsLI XIANG, Yali Li, Yuan Wang, Shengjin WangCVPR 2026
- Unifying Perception and Action: A Hybrid-Modality Pipeline with Implicit Visual Chain-of-Thought for Robotic Action GenerationXiangkai Ma, Lekai Xing, Han Zhang, Wenzhong Li 等CVPR 2026 · 被引用 11 次
- HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action ModelsMinghui Lin, Pengxiang Ding, Shu Wang, Zifeng Zhuang 等CVPR 2026 · 被引用 35 次
- Evo-1: Lightweight Vision-Language-Action Model with Preserved Semantic AlignmentTao Lin, Yilei Zhong, Yuxin Du, Jingjing Zhang 等CVPR 2026 · 被引用 47 次
- On Robustness of Vision-Language-Action Model against Multi-Modal PerturbationsJianing Guo, Zhenhong Wu, Chang Tu, Yiyao Ma 等ICLR 2026 · 被引用 7 次
