Lune

CVPR2026Top-tier venue

MonoVLM: Monocular 3D Visual Grounding with Vision Language Models

Huaizhi Qu, Hossein Nourkhiz Mahjoub, Vaishnav Tadiparthi, Kwonjoon Lee, Tianlong Chen

2026Year

Abstract

Figure 1. We propose MonoVLM, a simple yet effective method to equip Vision-Language Models (VLMs) with robust monocular 3D grounding capabilities. (a) The model takes an image and the textual query to predict the 3D bounding box (GT and prediction). (b) Even the latest large-scale VLMs struggle to understand 3D structure from 2D images. Our resulting model not only achieves significantly better results than these VLMs but also surpasses specialized vision-only models designed for this task.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 366772fc-6299-40b0-9fa4-91f49562e306

Builds on14

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines