Code-as-Monitor: Constraint-aware Visual Programming for Reactive and Proactive Robotic Failure Detection
Enshen Zhou, Qi Su, Cheng Chi, Zhizheng Zhang, Zhongyuan Wang, Tiejun Huang, Lu Sheng, He Wang
Abstract
Proactive Failure Detection (a) Reactive Failure Detection Task: Move the pan with the lobster to the stove, and be careful not to let the lobster drop out. (c) Real-world Test Pan Lobster Stove jump out Figure 1. For the task "Move the pan with lobster to the stove without losing the lobster", (a) reactive failure detection identifies failures after they occur, and (b) proactive failure detection prevents foreseeable failures. In (a), at t R 4 , the robot detects the failure after the lobster unpredictably jumps out due to the heat. In (b), pan tilting is detected at t P 3 and corrected it at t P ′ 3 , requiring real-time precision. We formulate both tasks as spatio-temporal constraint satisfaction problems, leveraging our proposed constraint elements for precise, real-time checking. For example, in (a), a large relative distance between pan and lobster indicates failure; in (b), a large angle between the pan and the horizontal plane needs correction. (c) shows that our method combined with an open-loop policy forms a closed-loop system, enabling proactive (e.g., detecting moving glass during grasping) and reactive (e.g., removing toy after grasping) failure detection in cluttered scenes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eda7c7c8-1b4f-49e0-acc8-91e3d9e55b34Cited by top-tier papers16
- RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for RoboticsEnshen Zhou, Jingkun An, Cheng Chi, Yi Han et al.NeurIPS 2025 · 159 citations
- SoFar: Language-Grounded Orientation Bridges Spatial Reasoning and Object ManipulationZekun Qi, Wenyao Zhang, Yufei Ding, Runpei Dong et al.NeurIPS 2025 · 65 citations
- Reason-RFT: Reinforcement Fine-Tuning for Visual Reasoning of Vision Language ModelsHuajie Tan, Yuheng Ji, Xiaoshuai Hao, Xiansheng Chen et al.NeurIPS 2025 · 45 citations
- From Language to Locomotion: Retargeting-free Humanoid Control via Motion Latent GuidanceZhe Li, Yangyang Wei, Boan Zhu, Yibo Peng et al.ICLR 2026 · 29 citations
- ManipLVM-R1: Reinforcement Learning for Reasoning in Embodied Manipulation with Large Vision-Language ModelsZirui Song, Guangxian Ouyang, Mingzhe Li, Yuheng Ji et al.AAAI 2026 · 21 citations
Builds on19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMsSoroush Nasiriany, Fei Xia, Wenhao Yu, Ted Xiao et al.ICML 2024 · 212 citations
- Generalized Planning in PDDL Domains with Pretrained Large Language ModelsTom Silver, Soham Dan, Kavitha Srinivas, Joshua B. Tenenbaum et al.AAAI 2024 · 194 citations
Related papers
- SAFE: Multitask Failure Detection for Vision-Language-Action ModelsQiao Gu, Yuanliang Ju, Shengxiang Sun, Igor Gilitschenski et al.NeurIPS 2025 · 103 citations
- RoboFailRing: Retrieval-Augmented and Language Grounding Failure Detection for VLM-enabled Robotic ManipulationChenduo Ying, Linkang Du, Yuanchao Shu, Peng ChengACL 2026
- Self-Correcting Robot Manipulation via Gaussian-Splatted ForesightShaohui Pan, Yong Xu, Ruotao Xu, Zihan Zhou et al.AAAI 2025 · 2 citations
- Self-Refining Vision Language Model for Robotic Failure Detection and ReasoningCarl Qi, Xiaojie Wang, Silong Yong, Stephen Sheng et al.ICLR 2026 · 8 citations
- PREGO: Online Mistake Detection in PRocedural EGOcentric VideosAlessandro Flaborea, Guido Maria D'Amely di Melendugno, Leonardo Plini, Luca Scofano et al.CVPR 2024
