Visual Grounding for Object Questions
Martin Nicolas Everaert, Xiruo Liu, Hiroyuki Takeda, Raja Bala, Vivek Yadav, Vidya Narayanan
Abstract
Current visual grounding research remains limited for practical applications, because existing techniques primarily focus on direct visual queries (e.g., "find the red car") or reading visible text (e.g., "what is the title of this book?"), rather than supporting general questions about objects (e.g., "how comfortable are these earbuds?"). We introduce the novel problem of Visual Grounding for Object Questions (VGOQ). Unlike previous work that grounds only what is directly visible in images, VGOQ handles open-ended general questions about objects, including concepts such as ease and comfort of use, and aims to identify visual evidence or context that would support an answer. This unexplored problem has immediate practical value, particularly in designing and optimizing product imagery in e-commerce stores. As initial steps toward this challenging task, we develop two automated data generation techniques combining existing models and data, and create two new datasets: ABO-VGOQ and VizWiz-VGOQ.We show that the data can be used to train a lightweight visual grounding model, and evaluate it against state-of-the-art approaches. Our results provide initial evidence that VGOQ represents a meaningful research direction: current SOTA visual grounding performance decreases from 29.2%-52.2% gIoU to 22.6%-37.2% gIoU when questions are rephrased from visual questions (segmentation of the answer) to general object questions (VizWiz-VGOQ, segmentation of visual evidence). On our new ABO-VGOQ dataset, our lightweight model achieves 39.5% gIoU, while current SOTA visual grounding approaches achieve only 12.4%-19.3%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning FrameworkPeng Wang, An Yang, Rui Men, Junyang Lin et al.ICML 2022 · 1,058 citations
- Image Segmentation Using Text and Image PromptsTimo Lüddecke, Alexander S. EckerCVPR 2022 · 457 citations
- ABO: Dataset and Benchmarks for Real-World 3D Object UnderstandingJasmine Collins, Shubham Goel, Kenan Deng, Achleshwar Luthra et al.CVPR 2022 · 117 citations
- GLaMM: Pixel Grounding Large Multimodal ModelHanoona Abdul Rasheed, Muhammad Maaz, Sahal Shaji Mullappilly, Abdelrahman M. Shaker et al.CVPR 2024 · 113 citations
Related papers
- Grounding Answers for Visual Questions Asked by Visually Impaired PeopleChongyan Chen, Samreen Anjum, Danna GurariCVPR 2022 · 48 citations
- Location-Aware Visual Question Generation with Lightweight ModelsNicholas Collin Suwono, Justin Chih-Yao Chen, Tun-Min Hung, Ting-Hao 'Kenneth' Huang et al.EMNLP 2023 · 3 citations
- Omni-Q: Omni-Directional Scene Understanding for Unsupervised Visual GroundingSai Wang, Yutian Lin, Yu WuCVPR 2024 · 1 citation
- VQA Therapy: Exploring Answer Differences by Visually Grounding AnswersChongyan Chen, Samreen Anjum, Danna GurariICCV 2023 · 20 citations
- Visually Precise QueryRiddhiman Dasgupta, Francis Tom, Sudhir Kumar, Mithun Das Gupta et al.ACM MM 2020 · 1 citation
