Intent3D: 3D Object Detection in RGB-D Scans Based on Human Intention
Weitai Kang, Mengxue Qu, Jyoti Kini, Yunchao Wei, Mubarak Shah, Yan Yan
Abstract
In real-life scenarios, humans seek out objects in the 3D world to fulfill their daily needs or intentions. This inspires us to introduce 3D intention grounding, a new task in 3D object detection employing RGB-D, based on human intention, such as "I want something to support my back." Closely related, 3D visual grounding focuses on understanding human reference. To achieve detection based on human intention, it relies on humans to observe the scene, reason out the target that aligns with their intention ("pillow" in this case), and finally provide a reference to the AI system, such as "A pillow on the couch". Instead, 3D intention grounding challenges AI agents to automatically observe, reason and detect the desired target solely based on human intention. To tackle this challenge, we introduce the new Intent3D dataset, consisting of 44,990 intention texts associated with 209 fine-grained classes from 1,042 scenes of the ScanNet [Dai et al., 2017] dataset. We also establish several baselines based on different language-based 3D object detection models on our benchmark. Finally, we propose IntentNet, our unique approach, designed to tackle this intention-based detection problem. It focuses on three key aspects: intention understanding, reasoning to identify object candidates, and cascaded adaptive learning that leverages the intrinsic priority logic of different losses for multiple objective optimization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- GPT4Scene: Understand 3D Scenes from Videos with Vision-Language ModelsZhangyang Qi, Zhixiong Zhang, Ye Fang, Jiaqi Wang et al.ICLR 2026 · 121 citations
- MLLM-For3D: Adapting Multimodal Large Language Model for 3D Reasoning SegmentationJiaxin Huang, Runnan Chen, Ziwen Li, Zhengqing Gao et al.NeurIPS 2025 · 18 citations
- GuirlVG: Incentivize GUI Visual Grounding via Empirical Exploration on Reinforcement LearningWeitai Kang, Bin Lei, Gaowen Liu, Caiwen Ding et al.ICLR 2026 · 6 citations
- Robin3D Improving 3D Large Language Model via Robust Instruction TuningWeitai Kang, Haifeng Huang, Yuzhang Shang, Mubarak Shah et al.ICCV 2025 · 6 citations
- VGent: Visual Grounding via Modular Design for Disentangling Reasoning and PredictionWeitai Kang, Jason Kuen, Mengwei Ren, Zijun Wei et al.CVPR 2026 · 5 citations
Builds on17
- 3D-LLM: Injecting the 3D World into Large Language ModelsYining Hong, Haoyu Zhen, Peihao Chen, Shuhong Zheng et al.NeurIPS 2023 · 662 citations
- An End-to-End Transformer Model for 3D Object DetectionIshan Misra, Rohit Girdhar, Armand JoulinICCV 2021 · 602 citations
- Ferret: Refer and Ground Anything Anywhere at Any GranularityHaoxuan You, Haotian Zhang, Zhe Gan, Xianzhi Du et al.ICLR 2024 · 515 citations
- Group-Free 3D Object Detection via TransformersZe Liu, Zheng Zhang, Yue Cao, Han Hu et al.ICCV 2021 · 368 citations
- 3D-VisTA: Pre-trained Transformer for 3D Vision and Text AlignmentZiyu Zhu, Xiaojian Ma, Yixin Chen, Zhidong Deng et al.ICCV 2023 · 247 citations
Related papers
- Visual Intention Grounding for Egocentric AssistantsPengzhan Sun, Junbin Xiao, Tze Ho Elden Tse, Yicong Li et al.ICCV 2025 · 2 citations
- Grounding 3D Object Affordance from 2D Interactions in ImagesYuhang Yang, Wei Zhai, Hongchen Luo, Yang Cao et al.ICCV 2023 · 69 citations
- ScanQA: 3D Question Answering for Spatial Scene UnderstandingDaichi Azuma, Taiki Miyanishi, Shuhei Kurita, Motoaki KawanabeCVPR 2022 · 135 citations
- Grounding 3D Object Affordance with Language Instructions, Visual Observations and InteractionsHe Zhu, Quyu Kong, Kechun Xu, Xunlong Xia et al.CVPR 2025
- ViGiL3D: A Linguistically Diverse Dataset for 3D Visual GroundingAustin T. Wang, ZeMing Gong, Angel X. ChangACL 2025 · 6 citations
