Beyond Human Perception: Understanding Multi-Object World from Monocular View
Keyu Guo, Yongle Huang, Shijie Sun, Xiangyu Song, Mingtao Feng, Zedong Liu, Huansheng Song, Tiantian Wang, Jianxin Li, Naveed Akhtar, Ajmal Saeed Mian
Abstract
Language and binocular vision play a crucial role in human understanding of the world. Advancements in artificial intelligence have also made it possible for machines to develop 3D perception capabilities essential for high-level scene understanding. However, only monocular cameras are often available in practice due to cost and space constraints. Enabling machines to achieve accurate 3D understanding from a monocular view is practical but presents significant challenges. We introduce MonoMulti-3DVG, a novel task aimed at achieving multi-object 3D Visual Grounding (3DVG) based on monocular RGB images, allowing machines to better understand and interact with the 3D world. To this end, we construct a large-scale benchmark dataset, MonoMulti3D-ROPE, and propose a model, CyclopsNet that integrates a State-Prompt Visual Encoder (SPVE) module with a Denoising Alignment Fusion (DAF) module to achieve robust multi-modal semantic alignment and fusion. This leads to more stable and robust multi-modal joint representations for downstream tasks. Experimental results show that our method significantly outperforms existing techniques on the MonoMulti3D-ROPE dataset. Our dataset and code are available at https://github.com/JasonHuang516/MonoMulti-3DVG
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 00619e23-11b2-4f24-b0d6-1d962479cdcbBuilds on14
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
- 3DVG-Transformer: Relation Modeling for Visual Grounding on Point CloudsLichen Zhao, Daigang Cai, Lu Sheng, Dong XuICCV 2021 · 234 citations
- MonoDTR: Monocular 3D Object Detection with Depth-Aware TransformerKuan-Chih Huang, Tsung-Han Wu, Hung-Ting Su, Winston H. HsuCVPR 2022 · 199 citations
- InstanceRefer: Cooperative Holistic Understanding for Visual Grounding on Point Clouds through Instance Multi-level Contextual ReferringZhihao Yuan, Xu Yan, Yinghong Liao, Ruimao Zhang et al.ICCV 2021 · 188 citations
- SAT: 2D Semantics Assisted Training for 3D Visual GroundingZhengyuan Yang, Songyang Zhang, Liwei Wang, Jiebo LuoICCV 2021 · 166 citations
Related papers
- Mono3DVG: 3D Visual Grounding in Monocular ImagesYang Zhan, Yuan Yuan, Zhitong XiongAAAI 2024 · 38 citations
- Mono3DVG-EnSD: Enhanced Spatial-aware and Dimension-decoupled Text Encoding for Monocular 3D Visual GroundingYuzhen Li, Min Liu, Zhaoyang Li, Yuan Bian et al.AAAI 2026
- MonoVLM: Monocular 3D Visual Grounding with Vision Language ModelsHuaizhi Qu, Hossein Nourkhiz Mahjoub, Vaishnav Tadiparthi, Kwonjoon Lee et al.CVPR 2026
- Unifying 2D and 3D Vision-Language UnderstandingAyush Jain, Alexander Swerdlow, Yuzhou Wang, Sergio Arnaud et al.ICML 2025
- Dual Enhancement on 3D Vision-Language Perception for Monocular 3D Visual GroundingYuzhen Li, Min Liu, Yuan Bian, Xueping Wang et al.ACM MM 2025 · 1 citation
