Robotic Visual Instruction
Yanbang Li, Ziyang Gong, Haoyang Li, Xiaoqi Huang, Haolan Kang, Guangping Bai, Xianzheng Ma
2025Year
9Top-tier citations
Abstract
Figure 1. (Left) Robotic visual instruction is a hand-drawn approach for commanding robots, utilizing circles and arrows to convey task definition. In long-horizon tasks, green and blue sketches denote the first and second task steps, respectively. (Right) It illustrates the action sequences output via VIEW. Our method exhibits robust generalization to real-world manipulation tasks, including (a) trajectory-following tasks, (b) cluttered environments with disturbances, and (c) multi-step operations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- PixelVLA: Advancing Pixel-level Understanding in Vision-Language-Action ModelWenqi Liang, Gan Sun, Yao He, Jiahua Dong et al.ICLR 2026 · 20 citations
- MLLM-For3D: Adapting Multimodal Large Language Model for 3D Reasoning SegmentationJiaxin Huang, Runnan Chen, Ziwen Li, Zhengqing Gao et al.NeurIPS 2025 · 18 citations
- Diagnose, Correct, and Learn from Manipulation Failures via Visual SymbolsXianchao Zeng, Xinyu Zhou, Youcheng Li, Jiayou Shi et al.CVPR 2026 · 18 citations
- From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot ManipulationYajie Li, Bozhou Zhang, Chun Gu, Zipei Ma et al.ICML 2026 · 2 citations
- VPN: Visual Prompt NavigationShuo Feng, Zihan Wang, Yuchen Li, Rui Kong et al.AAAI 2026 · 2 citations
Builds on10
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch et al.ICML 2023 · 2,601 citations
- Scaling Open-Vocabulary Object DetectionMatthias Minderer, Alexey A. Gritsenko, Neil HoulsbyNeurIPS 2023 · 482 citations
Related papers
- Object-Centric Prompt-Driven Vision-Language-Action Model for Robotic ManipulationXiaoqi Li, Jingyun Xu, Mingxu Zhang, Jiaming Liu et al.CVPR 2025
- RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory SketchesJiayuan Gu, Sean Kirmani, Paul Wohlhart, Yao Lu et al.ICLR 2024 · 135 citations
- VIP: Vision Instructed Pre-training for Robotic ManipulationZhuoling Li, Liangliang Ren, Jinrong Yang, Yong Zhao et al.ICML 2025
- InstructFlow: Adaptive Symbolic Constraint-Guided Code Generation for Long-Horizon PlanningHaotian Chi, Zeyu Feng, Yueming Lyu, Chengqi Zheng et al.NeurIPS 2025 · 6 citations
- ManipulaTHOR: A Framework for Visual Object ManipulationKiana Ehsani, Winson Han, Alvaro Herrasti, Eli VanderBilt et al.CVPR 2021
