Extending Phrase Grounding with Pronouns in Visual Dialogues
Panzhong Lu, Xin Zhang, Meishan Zhang, Min Zhang
Abstract
Conventional phrase grounding aims to localize noun phrases mentioned in a given caption to their corresponding image regions, which has achieved great success recently. Apparently, sole noun phrase grounding is not enough for cross-modal visual language understanding. Here we extend the task by considering pronouns as well. First, we construct a dataset of phrase grounding with both noun phrases and pronouns to image regions. Based on the dataset, we test the performance of phrase grounding by using a state-of-the-art literature model of this line. Then, we enhance the baseline grounding model with coreference information which should help our task potentially, modeling the coreference structures with graph convolutional networks. Experiments on our dataset, interestingly, show that pronouns are easier to ground than noun phrases, where the possible reason might be that these pronouns are much less ambiguous. Additionally, our final model with coreference information can significantly boost the grounding performance of both noun phrases and pronouns.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fed79241-6199-494d-a8bd-6d15cb3ef0c7Cited by top-tier papers3
- Target-Guided Composed Image RetrievalHaokun Wen, Xian Zhang, Xuemeng Song, Yinwei Wei et al.ACM MM 2023 · 53 citations
- Grounded Entity-Landmark Adaptive Pre-training for Vision-and-Language NavigationYibo Cui, Liang Xie, Yakun Zhang, Meishan Zhang et al.ICCV 2023 · 31 citations
- Disambiguating Reference in Visually Grounded Dialogues through Joint Modeling of Textual and Multimodal Semantic StructuresShun Inadumi, Nobuhiro Ueda, Koichiro YoshinoACL 2025
Builds on14
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve et al.ICCV 2021 · 1,114 citations
- TransVG: End-to-End Visual Grounding with TransformersJiajun Deng, Zhengyuan Yang, Tianlang Chen, Wengang Zhou et al.ICCV 2021 · 468 citations
- TACo: Token-aware Cascade Contrastive Learning for Video-Text AlignmentJianwei Yang, Yonatan Bisk, Jianfeng GaoICCV 2021 · 159 citations
- CorefQA: Coreference Resolution as Query-based Span PredictionWei Wu, Fei Wang, Arianna Yuan, Fei Wu et al.ACL 2020 · 153 citations
- Learning Cross-Modal Context Graph for Visual GroundingYongfei Liu, Bo Wan, Xiaodan Zhu, Xuming HeAAAI 2020 · 100 citations
Related papers
- Visual-Semantic Graph Matching for Visual GroundingChenchen Jing, Yuwei Wu, Mingtao Pei, Yao Hu et al.ACM MM 2020 · 35 citations
- Cross-Modal Omni Interaction Modeling for Phrase GroundingTianyu Yu, Tianrui Hui, Zhihao Yu, Yue Liao et al.ACM MM 2020 · 14 citations
- G3raphGround: Graph-Based Language GroundingMohit Bajaj, Lanjun Wang, Leonid SigalICCV 2019 · 67 citations
- Improving Zero-Shot Phrase Grounding via Reasoning on External Knowledge and Spatial RelationsZhan Shi, Yilin Shen, Hongxia Jin, Xiaodan ZhuAAAI 2022 · 7 citations
- Look Around Before Locating: Considering Content and Structure Information for Visual GroundingShiyi Zheng, Peizhi Zhao, Zhilong Zheng, Peihang He et al.AAAI 2025 · 3 citations
