Action-Sketcher: From Reasoning to Action via Visual Sketches for Robotic Manipulation
Huajie Tan, Peterson Co, Yijie Xu, Shanyu Rong, Yuheng Ji, Cheng Chi, Xiansheng Chen, Zhongxia Zhao, Pengwei Wang, Zhongyuan Wang, Shanghang Zhang
摘要
Temporal Reasoning: The high-level instruction is "Clean the objects on the table". Okay, let's analyze step by step and break down tasks to accomplish this. The scene contains … What I have done is Nothing. Now I need to: "Pick up the cola bottle with the left arm and place into left basket". … … Sketching Spatial Reasoning: The cola bottle … in front of a green basket on the left. … The bounding box [42, 76, 74, 140]clearly frames the cola bottle. … The arrow originates at [59, 105], the center at [53, 28], and extends to [61, 77]… "bbox": [42, 76, 74, 140], "arrow": [[59, 105], [53, 28], [61, 77]]
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper34
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMsPeter Tong, Ellis Brown, Penghao Wu, Sanghyun Woo 等NeurIPS 2024 · 被引用 1,004 次
- Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction TuningFuxiao Liu, Kevin Lin, Linjie Li, Jianfeng Wang 等ICLR 2024 · 被引用 476 次
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic ManipulationTianxing Chen, Zanxin Chen, Baijun Chen, Zijian Cai 等ICML 2026 · 被引用 394 次
相关 Paper
- Visual Room RearrangementLuca Weihs, Matt Deitke, Aniruddha Kembhavi, Roozbeh MottaghiCVPR 2021
- Spatio-Temporal Graph for Video Captioning With Knowledge DistillationBoxiao Pan, Haoye Cai, De-An Huang, Kuan-Hui Lee 等CVPR 2020
- MesaTask: Towards Task-Driven Tabletop Scene Generation via 3D Spatial ReasoningJinkun Hao, Naifu Liang, Zhen Luo, Xudong Xu 等NeurIPS 2025 · 被引用 21 次
- Language-driven Grasp DetectionVuong Dinh An, Minh Nhat Vu, Baoru Huang, Nghia Nguyen 等CVPR 2024
- ChainHOI: Joint-based Kinematic Chain Modeling for Human-Object Interaction GenerationLing-An Zeng, Guohong Huang, Yi-Lin Wei, Shengbo Gu 等CVPR 2025
