CityVG: Contrastive Fine-Tuning and Reward-Based Chain-of-Thought Reasoning for Zero-Shot City-Scale 3D Visual Grounding
Jianjun Zhang, Hanli Wang
Abstract
3D Visual Grounding (3DVG) locates objects in 3D scenes based on natural language descriptions. However, existing methods are primarily confined to small-scale indoor data or rely on heavy supervision, failing to generalize to complex large-scale urban environments. To address this limitation, we present CityVG, the first city-scale zero-shot 3D visual grounding framework capable of localizing urban objects without manual annotations. Our approach adopts a retrieval-and-reasoning paradigm comprising two key components. Specifically, we propose a contrastive fine-tuning strategy to align textual queries with urban scene graphs. By leveraging an LLM-driven graph clustering mechanism, we automatically construct high-quality positive and negative training pairs and fine-tune the text encoder via contrastive learning, resulting in a scene-adaptive text encoder that enables efficient alignment without grounding supervision. Complementing this, a multi-trajectory reward-based Chain-of-Thought (CoT) reasoning strategy is designed for inference. This mechanism iteratively evaluates candidate objects by aggregating reward scores across diverse reasoning trajectories, selecting the target that is most consistent with both appearance and spatial constraints. Extensive experiments on city-scale 3D grounding benchmarks demonstrate that CityVG achieves strong zero-shot localization performance and generalizes effectively to unseen urban environments. The source code of this work can be found in https://mic.tongji.edu.cn.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on21
- LERF: Language Embedded Radiance FieldsJustin Kerr, Chung Min Kim, Ken Goldberg, Angjoo Kanazawa et al.ICCV 2023 · 620 citations
- 3D Scene Graph: A Structure for Unified Semantics, 3D Space, and CameraIro Armeni, Zhi-Yang He, Amir Zamir, JunYoung Gwak et al.ICCV 2019 · 474 citations
- 3D-VisTA: Pre-trained Transformer for 3D Vision and Text AlignmentZiyu Zhu, Xiaojian Ma, Yixin Chen, Zhidong Deng et al.ICCV 2023 · 247 citations
- 3DVG-Transformer: Relation Modeling for Visual Grounding on Point CloudsLichen Zhao, Daigang Cai, Lu Sheng, Dong XuICCV 2021 · 234 citations
- InstanceRefer: Cooperative Holistic Understanding for Visual Grounding on Point Clouds through Instance Multi-level Contextual ReferringZhihao Yuan, Xu Yan, Yinghong Liao, Ruimao Zhang et al.ICCV 2021 · 188 citations
Related papers
- Visual Programming for Zero-Shot Open-Vocabulary 3D Visual GroundingZhihao Yuan, Jinke Ren, Chun-Mei Feng, Hengshuang Zhao et al.CVPR 2024 · 19 citations
- UZ3DVG: Unaided Zero-Shot 3D Visual Grounding with Generated Language ConditionsWenbin Tan, Jiawen Lin, Yuan Xie, Yachao Zhang et al.CVPR 2026
- SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual GroundingJiawen Lin, Shiran Bian, Yihang Zhu, Wenbin Tan et al.ACM MM 2025 · 4 citations
- View-on-Graph: Zero-Shot 3D Visual Grounding via Vision-Language Reasoning on Scene GraphsYuanyuan Liu, Haiyang Mei, Dongyang Zhan, Jiayue Zhao et al.AAAI 2026 · 1 citation
- OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal GroundingMinghang Zheng, Zihao Yin, Yi Yang, Yuxin Peng et al.CVPR 2026 · 4 citations
