Language-driven Grasp Detection
Vuong Dinh An, Minh Nhat Vu, Baoru Huang, Nghia Nguyen, Hieu Le, Thieu Vo, Anh Nguyen
Abstract
A coffee cup, a grape biscuit, a black oven plate on a white dining table Grasp the coffee cup. Grasp the calculator at its keypad. A black pen, a steel marble and a digital calculator on an office desk A black pot, a green bottle, a bread container arranged on a kitchen counter Grab the neck of the green bottle. Give me the brown-band wristwatch. A brown-band wristwatch, a sleek black pen and an office clip placed on a table A white mug, a spiral notepad, a fountain pen resting on a brown desk Pick the fountain pen at its cap. Figure 1 . We present a new dataset and method for language-driven grasp task.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 727a7795-9304-41f8-ac43-cc710d871548Cited by top-tier papers6
- Bring My Cup! Personalizing Vision-Language-Action Models with Visual Attentive PromptingSangoh Lee, Sangwoo Mo, Wook-Shin HanICML 2026 · 5 citations
- RAGNet: Large-Scale Reasoning-Based Affordance Segmentation Benchmark Towards General GraspingDongming Wu, Yanping Fu, Saike Huang, Yingfei Liu et al.ICCV 2025 · 2 citations
- AffordMatcher: Affordance Learning in 3D Scenes from Visual SignifiersNghia Vu, Tuong Do, Khang Nguyen, Baoru Huang et al.CVPR 2026 · 2 citations
- RealVLG-R1: A Large-Scale Real-World Visual-Language Grounding Benchmark for Robotic Perception and ManipulationLinfei Li, Lin Zhang, Ying ShenCVPR 2026
- PhysVLM: Enabling Visual Language Models to Understand Robotic Physical ReachabilityWeijie Zhou, Manli Tao, Chaoyang Zhao, Haiyun Guo et al.CVPR 2025
Builds on33
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- Action-Sketcher: From Reasoning to Action via Visual Sketches for Robotic ManipulationHuajie Tan, Peterson Co, Yijie Xu, Shanyu Rong et al.CVPR 2026
- RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to ConcreteYuheng Ji, Huajie Tan, Jiayu Shi, Xiaoshuai Hao et al.CVPR 2025
- DexFuncGrasp: A Robotic Dexterous Functional Grasp Dataset Constructed from a Cost-Effective Real-Simulation Annotation SystemJinglue Hang, Xiangbo Lin, Tianqiang Zhu, Xuanheng Li et al.AAAI 2024 · 17 citations
- AffordDexGrasp: Open-Set Language-Guided Dexterous Grasp With Generalizable-Instructive AffordanceYi-Lin Wei, Mu Lin, Yuhao Lin, Jian-Jian Jiang et al.ICCV 2025 · 8 citations
- Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video ContentQiuheng Wang, Yukai Shi, Jiarong Ou, Rui Chen et al.CVPR 2025
