When Visual Grounding Meets Gigapixel-Level Large-Scale Scenes: Benchmark and Approach
M. Tao, Bing Bai, Haozhe Lin, Heyuan Wang, Yu Wang, Lin Luo, Lu Fang
Abstract
Visual grounding refers to the process of associating natural language expressions with corresponding regions within an image. Existing benchmarks for visual grounding primarily operate within small-scale scenes with a few objects. Nevertheless, recent advances in imaging technology have enabled the acquisition of gigapixel-level images, providing high-resolution details in large-scale scenes containing numerous objects. To bridge this gap between imaging and computer vision benchmarks and make grounding more practically valuable, we introduce a novel dataset, named GigaGrounding, designed to challenge visual grounding models in gigapixel-level large-scale scenes. We extensively analyze and compare the dataset with existing benchmarks, demonstrating that GigaGrounding presents unique challenges such as large-scale scene understanding, gigapixel-level resolution, significant variations in object scales, and the “multi-hop expressions”. Furthermore, we introduced a simple yet effective grounding approach, which employs a “glance-to-zoom-in” paradigm and exhibits enhanced capabilities for addressing the GigaGrounding task. The dataset is available at www.gigavision.ai.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aa68e343-d139-425e-9148-e6ea9a8cb67fCited by top-tier papers4
- SparseFormer: Detecting Objects in HRW Shots via Sparse Vision TransformerWenxi Li, Yuchen Guo, Jilai Zheng, Haozhe Lin et al.ACM MM 2024 · 3 citations
- GigaMoE: Sparsity-Guided Mixture of Experts for Efficient Gigapixel Object DetectionXiang Li, Wenxi Li, Yuetong Wang, Chenyang Lyu et al.AAAI 2026 · 1 citation
- Referring Expression Comprehension for Small ObjectsKanoko Goto, Takumi Hirose, Mahiro Ukai, Shuhei Kurita et al.ICCV 2025
- Hybrid Reciprocal Transformer with Triplet Feature Alignment for Scene Graph GenerationJiawei Fu, Tiantian Zhang, Kai Chen, Qi DouCVPR 2025
Builds on13
- TransVG: End-to-End Visual Grounding with TransformersJiajun Deng, Zhengyuan Yang, Tianlang Chen, Wengang Zhou et al.ICCV 2021 · 468 citations
- A Fast and Accurate One-Stage Approach to Visual GroundingZhengyuan Yang, Boqing Gong, Liwei Wang, Wenbing Huang et al.ICCV 2019 · 441 citations
- Referring Transformer: A One-step Approach to Multi-task Visual GroundingMuchen Li, Leonid SigalNeurIPS 2021 · 270 citations
- Dynamic Graph Attention for Referring Expression ComprehensionSibei Yang, Guanbin Li, Yizhou YuICCV 2019 · 251 citations
- Zero-Shot Grounding of Objects From Natural Language QueriesArka Sadhu, Kan Chen, Ram NevatiaICCV 2019 · 176 citations
Related papers
- GroundingME: Exposing the Visual Grounding Gap in MLLMs through Multi-Dimensional EvaluationRang Li, Lei Li, Shuhuai Ren, Hao Tian et al.CVPR 2026 · 10 citations
- Grounding-IQA: Grounding Multimodal Language Model for Image Quality AssessmentZheng Chen, Xun Zhang, Wenbo Li, Renjing Pei et al.ICLR 2026 · 12 citations
- AerialVG: A Challenging Benchmark for Aerial Visual Grounding by Exploring Positional RelationsJunli Liu, Qizhi Chen, Zhigang Wang, Yiwen Tang et al.ICCV 2025 · 5 citations
- Visual Grounding in Remote Sensing ImagesYuxi Sun, Shanshan Feng, Xutao Li, Yunming Ye et al.ACM MM 2022 · 76 citations
- ViGiL3D: A Linguistically Diverse Dataset for 3D Visual GroundingAustin T. Wang, ZeMing Gong, Angel X. ChangACL 2025 · 6 citations
