Improving Weakly Supervised Visual Grounding by Contrastive Knowledge Distillation
Liwei Wang, Jing Huang, Yin Li, Kun Xu, Zhengyuan Yang, Dong Yu
摘要
Weakly supervised phrase grounding aims at learning region-phrase correspondences using only image-sentence pairs. A major challenge thus lies in the missing links between image regions and sentence phrases during training. To address this challenge, we leverage a generic object detector at training time, and propose a contrastive learning framework that accounts for both region-phrase and imagesentence matching. Our core innovation is the learning of a region-phrase score function, based on which an imagesentence score function is further constructed. Importantly, our region-phrase score function is learned by distilling from soft matching scores between the detected object names and candidate phrases within an image-sentence pair, while the image-sentence score function is supervised by ground-truth image-sentence pairs. The design of such score functions removes the need of object detection at test time, thereby significantly reducing the inference cost. Without bells and whistles, our approach achieves state-of-the-art results on visual phrase grounding, surpassing previous methods that require expensive object detectors at test time.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper28
- GroupViT: Semantic Segmentation Emerges from Text SupervisionJiarui Xu, Shalini De Mello, Sifei Liu, Wonmin Byeon 等CVPR 2022 · 被引用 398 次
- Multi-View Transformer for 3D Visual GroundingShijia Huang, Yilun Chen, Jiaya Jia, Liwei WangCVPR 2022 · 被引用 97 次
- TubeDETR: Spatio-Temporal Video Grounding with TransformersAntoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev 等CVPR 2022 · 被引用 87 次
- Pseudo-Q: Generating Pseudo Language Queries for Visual GroundingHaojun Jiang, Yuanze Lin, Dongchen Han, Shiji Song 等CVPR 2022 · 被引用 60 次
- Referring Image Segmentation Using Text SupervisionFang Liu, Yuhao Liu, Yuqiu Kong, Ke Xu 等ICCV 2023 · 被引用 52 次
它引用的顶会 Paper10
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- A Fast and Accurate One-Stage Approach to Visual GroundingZhengyuan Yang, Boqing Gong, Liwei Wang, Wenbing Huang 等ICCV 2019 · 被引用 441 次
- Align2Ground: Weakly Supervised Phrase Grounding Guided by Image-Caption AlignmentSamyak Datta, Karan Sikka, Anirban Roy, Karuna Ahuja 等ICCV 2019 · 被引用 113 次
- G3raphGround: Graph-Based Language GroundingMohit Bajaj, Lanjun Wang, Leonid SigalICCV 2019 · 被引用 67 次
- Compact Trilinear Interaction for Visual Question AnsweringTuong Do, Huy Tran, Thanh-Toan Do, Erman Tjiputra 等ICCV 2019 · 被引用 63 次
相关 Paper
- Contrastive Learning with Expectation-Maximization for Weakly Supervised Phrase GroundingKeqin Chen, Richong Zhang, Samuel Mensah, Yongyi MaoEMNLP 2022 · 被引用 3 次
- Phrase Localization Without Paired Training ExamplesJosiah Wang, Lucia SpeciaICCV 2019 · 被引用 51 次
- Similarity Maps for Self-Training Weakly-Supervised Phrase GroundingTal Shaharabany, Lior WolfCVPR 2023
- Detector-Free Weakly Supervised Grounding by SeparationAssaf Arbelle, Sivan Doveh, Amit Alfassy, Joseph Shtok 等ICCV 2021 · 被引用 31 次
- What is Where by Looking: Weakly-Supervised Open-World Phrase-Grounding without Text InputsTal Shaharabany, Yoad Tewel, Lior WolfNeurIPS 2022 · 被引用 26 次
