Phrase Localization Without Paired Training Examples
Josiah Wang, Lucia Specia
Abstract
Localizing phrases in images is an important part of image understanding and can be useful in many applications that require mappings between textual and visual information. Existing work attempts to learn these mappings from examples of phrase-image region correspondences (strong supervision) or from phrase-image pairs (weak supervision). We postulate that such paired annotations are unnecessary, and propose the first method for the phrase localization problem where neither training procedure nor paired, task-specific data is required. Our method is simple but effective: we use off-the-shelf approaches to detect objects, scenes and colours in images, and explore different approaches to measure semantic similarity between the categories of detected visual elements and words in phrases. Experiments on two well-known phrase localization datasets show that this approach surpasses all weakly supervised methods by a large margin and performs very competitively to strongly supervised methods, and can thus be considered a strong baseline to the task. The non-paired nature of our method makes it applicable to any domain and where no paired phrase localization annotation is available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d90f5659-9e4f-4157-a470-5c1bfa782942Cited by top-tier papers14
- Pseudo-Q: Generating Pseudo Language Queries for Visual GroundingHaojun Jiang, Yuanze Lin, Dongchen Han, Shiji Song et al.CVPR 2022 · 60 citations
- Referring Image Segmentation Using Text SupervisionFang Liu, Yuhao Liu, Yuqiu Kong, Ke Xu et al.ICCV 2023 · 52 citations
- Deconfounded Visual GroundingJianqiang Huang, Yu Qin, Jiaxin Qi, Qianru Sun et al.AAAI 2022 · 38 citations
- Detector-Free Weakly Supervised Grounding by SeparationAssaf Arbelle, Sivan Doveh, Amit Alfassy, Joseph Shtok et al.ICCV 2021 · 31 citations
- Look at What I'm Doing: Self-Supervised Spatial Grounding of Narrations in Instructional VideosReuben Tan, Bryan A. Plummer, Kate Saenko, Hailin Jin et al.NeurIPS 2021 · 30 citations
Related papers
- Improving Weakly Supervised Visual Grounding by Contrastive Knowledge DistillationLiwei Wang, Jing Huang, Yin Li, Kun Xu et al.CVPR 2021
- What is Where by Looking: Weakly-Supervised Open-World Phrase-Grounding without Text InputsTal Shaharabany, Yoad Tewel, Lior WolfNeurIPS 2022 · 26 citations
- Similarity Maps for Self-Training Weakly-Supervised Phrase GroundingTal Shaharabany, Lior WolfCVPR 2023
- Cap2Det: Learning to Amplify Weak Caption Supervision for Object DetectionKeren Ye, Mingda Zhang, Adriana Kovashka, Wei Li et al.ICCV 2019 · 61 citations
- Shatter and Gather: Learning Referring Image Segmentation with Text SupervisionDongwon Kim, Namyup Kim, Cuiling Lan, Suha KwakICCV 2023 · 29 citations
