MonoVLM: Monocular 3D Visual Grounding with Vision Language Models
Huaizhi Qu, Hossein Nourkhiz Mahjoub, Vaishnav Tadiparthi, Kwonjoon Lee, Tianlong Chen
Abstract
Figure 1. We propose MonoVLM, a simple yet effective method to equip Vision-Language Models (VLMs) with robust monocular 3D grounding capabilities. (a) The model takes an image and the textual query to predict the 3D bounding box (GT and prediction). (b) Even the latest large-scale VLMs struggle to understand 3D structure from 2D images. Our resulting model not only achieves significantly better results than these VLMs but also surpasses specialized vision-only models designed for this task.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 366772fc-6299-40b0-9fa4-91f49562e306Builds on14
- MonoDETR: Depth-guided Transformer for Monocular 3D Object DetectionRenrui Zhang, Han Qiu, Tai Wang, Ziyu Guo et al.ICCV 2023 · 175 citations
- VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D ReconstructionZhiwen Fan, Jian Zhang, Renjie Li, Junge Zhang et al.CVPR 2026 · 171 citations
- Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry PriorsDuo Zheng, Shijia Huang, Yanyang Li, Liwei WangNeurIPS 2025 · 130 citations
- SpatialLM: Training Large Language Models for Structured Indoor ModelingYongsen Mao, Junhao Zhong, Chuan Fang, Jia Zheng et al.NeurIPS 2025 · 89 citations
- VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action ModelYihao Wang, Pengxiang Ding, Lingxiao Li, Can Cui et al.AAAI 2026 · 76 citations
Related papers
- Grounded 3D-Aware Spatial Vision-Language ModelingAn-Chieh Cheng, Yang Fu, Yatai Ji, Ligeng Zhu et al.CVPR 2026
- LLaVA³: Representing 3D Scenes Like a Cubist Painter to Boost 3D Scene Understanding of VLMsDoriand Petit, Steve Bourgeois, Vincent Gay-Bellile, Florian Chabot et al.AAAI 2026
- InteractVLM: 3D Interaction Reasoning from 2D Foundational ModelsSai Kumar Dwivedi, Dimitrije Antic, Shashank Tripathi, Omid Taheri et al.CVPR 2025
- LocateAnything3D: Vision-Language 3D Detection with Chain-of-SightYunze Man, Shihao Wang, Guowen Zhang, Johan Bjorck et al.CVPR 2026 · 6 citations
- SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding CapabilityJiankang Wang, Zhihan Zhang, Zhihang Liu, Yang Li et al.AAAI 2026 · 20 citations
