Action-Sketcher: From Reasoning to Action via Visual Sketches for Robotic Manipulation
Huajie Tan, Peterson Co, Yijie Xu, Shanyu Rong, Yuheng Ji, Cheng Chi, Xiansheng Chen, Zhongxia Zhao, Pengwei Wang, Zhongyuan Wang, Shanghang Zhang
Abstract
Temporal Reasoning: The high-level instruction is "Clean the objects on the table". Okay, let's analyze step by step and break down tasks to accomplish this. The scene contains … What I have done is Nothing. Now I need to: "Pick up the cola bottle with the left arm and place into left basket". … … Sketching Spatial Reasoning: The cola bottle … in front of a green basket on the left. … The bounding box [42, 76, 74, 140]clearly frames the cola bottle. … The arrow originates at [59, 105], the center at [53, 28], and extends to [61, 77]… "bbox": [42, 76, 74, 140], "arrow": [[59, 105], [53, 28], [61, 77]]
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on34
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMsPeter Tong, Ellis Brown, Penghao Wu, Sanghyun Woo et al.NeurIPS 2024 · 1,004 citations
- Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction TuningFuxiao Liu, Kevin Lin, Linjie Li, Jianfeng Wang et al.ICLR 2024 · 476 citations
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic ManipulationTianxing Chen, Zanxin Chen, Baijun Chen, Zijian Cai et al.ICML 2026 · 394 citations
Related papers
- Visual Room RearrangementLuca Weihs, Matt Deitke, Aniruddha Kembhavi, Roozbeh MottaghiCVPR 2021
- Spatio-Temporal Graph for Video Captioning With Knowledge DistillationBoxiao Pan, Haoye Cai, De-An Huang, Kuan-Hui Lee et al.CVPR 2020
- MesaTask: Towards Task-Driven Tabletop Scene Generation via 3D Spatial ReasoningJinkun Hao, Naifu Liang, Zhen Luo, Xudong Xu et al.NeurIPS 2025 · 21 citations
- Language-driven Grasp DetectionVuong Dinh An, Minh Nhat Vu, Baoru Huang, Nghia Nguyen et al.CVPR 2024
- ChainHOI: Joint-based Kinematic Chain Modeling for Human-Object Interaction GenerationLing-An Zeng, Guohong Huang, Yi-Lin Wei, Shengbo Gu et al.CVPR 2025
