Improving Zero-Shot Phrase Grounding via Reasoning on External Knowledge and Spatial Relations
Zhan Shi, Yilin Shen, Hongxia Jin, Xiaodan Zhu
Abstract
Phrase grounding is a multi-modal problem that localizes a particular noun phrase in an image referred to by a text query. In the challenging zero-shot phrase grounding setting, the existing state-of-the-art grounding models have limited capacity in handling the unseen phrases. Humans, however, can ground novel types of objects in images with little effort, significantly benefiting from reasoning with commonsense. In this paper, we design a novel phrase grounding architecture that builds multi-modal knowledge graphs using external knowledge and then performs graph reasoning and spatial relation reasoning to localize the referred nouns phrases. We perform extensive experiments on different zero-shot grounding splits sub-sampled from the Flickr30K Entity and Visual Genome dataset, demonstrating that the proposed framework is orthogonal to backbone image encoders and outperforms the baselines by 2 3% in accuracy, resulting in a significant improvement under the standard evaluation metrics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4c1a5f0d-5f54-4979-83a6-b5769f77c12aCited by top-tier papers1
Ask how each one uses itBuilds on4
- A Fast and Accurate One-Stage Approach to Visual GroundingZhengyuan Yang, Boqing Gong, Liwei Wang, Wenbing Huang et al.ICCV 2019 · 441 citations
- Zero-Shot Grounding of Objects From Natural Language QueriesArka Sadhu, Kan Chen, Ram NevatiaICCV 2019 · 176 citations
- Boosting Visual Question Answering with Context-aware Knowledge AggregationGuohao Li, Xin Wang, Wenwu ZhuACM MM 2020 · 82 citations
- From Strings to Things: Knowledge-Enabled VQA Model That Can Read and ReasonAjeet Kumar Singh, Anand Mishra, Shashank Shekhar, Anirban ChakrabortyICCV 2019 · 54 citations
Related papers
- Learning Cross-Modal Context Graph for Visual GroundingYongfei Liu, Bo Wan, Xiaodan Zhu, Xuming HeAAAI 2020 · 100 citations
- G3raphGround: Graph-Based Language GroundingMohit Bajaj, Lanjun Wang, Leonid SigalICCV 2019 · 67 citations
- Extending Phrase Grounding with Pronouns in Visual DialoguesPanzhong Lu, Xin Zhang, Meishan Zhang, Min ZhangEMNLP 2022 · 5 citations
- Cross-Modal Omni Interaction Modeling for Phrase GroundingTianyu Yu, Tianrui Hui, Zhihao Yu, Yue Liao et al.ACM MM 2020 · 14 citations
- Triple Alignment Strategies for Zero-shot Phrase Grounding under Weak SupervisionPengyue Lin, Ruifan Li, Yuzhe Ji, Zhihan Yu et al.ACM MM 2024 · 2 citations
