Small Object, Great Challenge: A Benchmark for Small Object Visual Grounding
Wenqi Jia, Ruifan Li, Pengyue Lin, Fangxiang Feng, Zhanyu Ma, Xiaojie Wang
Abstract
The task of visual grounding (i.e., VG) aims to locate or segment objects in images based on referring expressions. Existing research on VG primarily focuses on large objects. However, these images often contain objects at various scales. Although large objects are usually the visual focus, small objects sometimes carry crucial information. To bridge the gap, we propose a novel benchmark for small object visual grounding, i.e., SoVG. Specifically, we introduce an automatic pipeline using MLLMs to build a benchmark dataset. Our pipeline is built on the popular dataset COCO. Thus, we obtain our RefCO-COs dataset. The visual objects in our RefCOCOs have an average area of 1/50 area of an entire image, whereas that of classic VG datasets is 1/5. Furthermore, we propose SoVG-Net with a hierarchical textual infusion module for the novel SoVG task. Finally, we conduct extensive experiments using classic datasets with our RefCO-COs. The results showcase that our built dataset is useful for advancing VG research, and our proposed SoVG-Net is a strong baseline. Our code and models are available at https://github.com/lemonskyer/sovg.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ccab51b6-0faa-447f-ae2f-aba047ecf5dbBuilds on19
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve et al.ICCV 2021 · 1,114 citations
- OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning FrameworkPeng Wang, An Yang, Rui Men, Junyang Lin et al.ICML 2022 · 1,058 citations
- TransVG: End-to-End Visual Grounding with TransformersJiajun Deng, Zhengyuan Yang, Tianlang Chen, Wengang Zhou et al.ICCV 2021 · 468 citations
- Learning to Assemble Neural Module Tree Networks for Visual GroundingDaqing Liu, Hanwang Zhang, Feng Wu, Zheng-Jun ZhaICCV 2019 · 317 citations
- Referring Transformer: A One-step Approach to Multi-task Visual GroundingMuchen Li, Leonid SigalNeurIPS 2021 · 270 citations
Related papers
- Unveiling Parts Beyond Objects: Towards Finer-Granularity Referring Expression SegmentationWenxuan Wang, Tongtian Yue, Yisi Zhang, Longteng Guo et al.CVPR 2024 · 5 citations
- Towards Further Comprehension on Referring Expression with RationaleRengang Li, Baoyu Fan, Xiaochuan Li, Runze Zhang et al.ACM MM 2022 · 2 citations
- Visual Grounding in Remote Sensing ImagesYuxi Sun, Shanshan Feng, Xutao Li, Yunming Ye et al.ACM MM 2022 · 76 citations
- Task-aware Cross-modal Feature Refinement Transformer with Large Language Models for Visual GroundingWenbo Chen, Zhen Xu, Ruotao Xu, Si Wu et al.CVPR 2025
- MC-Bench: A Benchmark for Multi-Context Visual Grounding in the Era of MLLMsYunqiu Xu, Linchao Zhu, Yi YangICCV 2025 · 7 citations
