Sentinel-VLA: A Metacognitive VLA Model with Active Status Monitoring for Dynamic Reasoning and Error Recovery
Wenhao Li, Xiu Su, Dan Niu, Yichao Cao, Hongyan Xu, Zhe Qu, Lei Fan, Shan You, Chang Xu
Abstract
Vision-language-action (VLA) models have advanced the field of embodied manipulation by harnessing broad world knowledge and strong generalization. However, current VLA models still face several key challenges, including limited reasoning capability, lack of status monitoring, and difficulty in self-correction. In this paper, we introduce Sentinel-VLA, a metacognitive VLA model equipped with an active ``sentinel'' module to monitor real-time execution status. Only when necessary, such as during initial planning or upon detecting an error, the model triggers a dynamic reasoning or formulate error recovery solutions. This on-demand reasoning mechanism ensures robust decision-making while minimizing computational overhead. Notably, all training data (spanning 44 tasks and over 2.6 million transitions) is automatically generated and annotated through our designed pipeline. We also propose the Self-Evolving Continual Learning (SECL) algorithm, which allows Sentinel-VLA to identify its capability boundaries and automatically collect data for expansion, paired with Orthogonal Continual Adapter (OC-Adapter) to constrain parameter updates to an orthogonal space, thereby preventing catastrophic forgetting. Real-world experiments demonstrate that Sentinel-VLA boosts the task success rate by over 30% compared to the SOTA model, PI0. We will open-source all the code, weights, and data generation pipeline.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f99d8d6d-ee9d-40bd-9683-d7ab32da29a7Cited by top-tier papers4
- Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Vision-Language ModelsChengcheng Wang, Jianyuan Guo, Hongguang Li, Yuchuan Tian et al.ICML 2026 · 14 citations
- VLA-ATTC: Adaptive Test-Time Compute for VLA Models with Relative Action Critic ModelWenhao Li, Xiu Su, Yichao Cao, Hongyan Xu et al.ICML 2026 · 13 citations
- See What Matters: Differentiable Grid Sample Pruning for Generalizable Vision-Language-Action ModelYixu Feng, Zinan Zhao, Yanxiang Ma, Chenghao Xia et al.ICML 2026 · 6 citations
- 4DPChat: Towards Dynamic Point Cloud Understanding with Failure-Aware BootstrappingXindan Zhang, Weilong Yan, YUFEI SHI, Xuerui Qiu et al.ICML 2026 · 6 citations
Builds on13
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- VLA-Cache: Efficient Vision-Language-Action Manipulation via Adaptive Token CachingSiyu Xu, Yunke Wang, Chenghao Xia, Dihao Zhu et al.NeurIPS 2025 · 95 citations
- Self-supervised Geometric Features Discovery via Interpretable Attention for Vehicle Re-Identification and BeyondMing Li, Xinming Huang, Ziming ZhangICCV 2021 · 55 citations
- Affordance Field Intervention: Enabling VLAs to Escape Memory Traps in Robotic ManipulationSiyu Xu, Zijian Wang, Yunke Wang, Chenghao Xia et al.CVPR 2026 · 14 citations
- Do You Have Freestyle? Expressive Humanoid Locomotion via Audio ControlZhe Li, Cheng Chi, Yangyang Wei, Boan Zhu et al.CVPR 2026 · 13 citations
Related papers
- OneTwoVLA: A Unified Vision-Language-Action Model with Adaptive ReasoningFanqi Lin, Ruiqian Nai, Yingdong Hu, Jiacheng You et al.ICLR 2026 · 129 citations
- NeurVLA: Unleashing Failure-Handling Capability of Vision-Language-Action Models via Neural-Symbolic ReasoningXuqi Liu, Minghe Gao, Juncheng Li, Siliang TangICML 2026
- SemanticVLA: Towards Semantic Reasoning over Action Memorization via Synergistic Explicit Trace and Latent Action PlanningFei Ni, Zhuo Chen, Yifu Yuan, Zibin Dong et al.CVPR 2026
- CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action ModelsQingqing Zhao, Yao Lu, Moo Jin Kim, Zipeng Fu et al.CVPR 2025
- Vision-Language-Action Instruction Tuning: From Understanding to ManipulationShuai Yang, Hao Li, Bin Wang, Yilun Chen et al.ICLR 2026 · 50 citations
