Improving Visual Grounding by Encouraging Consistent Gradient-Based Explanations
Ziyan Yang, Kushal Kafle, Franck Dernoncourt, Vicente Ordonez
摘要
We propose a margin-based loss for tuning joint visionlanguage models so that their gradient-based explanations are consistent with region-level annotations provided by humans for relatively smaller grounding datasets. We refer to this objective as Attention Mask Consistency (AMC) and demonstrate that it produces superior visual grounding results than previous methods that rely on using visionlanguage models to score the outputs of object detectors. Particularly, a model trained with AMC on top of standard vision-language modeling objectives obtains a state-of-theart accuracy of 86.49% in the Flickr30k visual grounding benchmark, an absolute improvement of 5.38% when compared to the best previous model trained under the same level of supervision. Our approach also performs exceedingly well on established benchmarks for referring expression comprehension where it obtains 80.34% accuracy in the easy test of RefCOCO+, and 64.55% in the difficult split. AMC is effective, easy to implement, and is general as it can be adopted by any vision-language model, and can use any type of region annotations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- ENCODER: Entity Mining and Modification Relation Binding for Composed Image RetrievalZixu Li, Zhiwei Chen, Haokun Wen, Zhiheng Fu 等AAAI 2025 · 被引用 59 次
- Studying How to Efficiently and Effectively Guide Models with ExplanationsSukrut Rao, Moritz Böhle, Amin Parchami-Araghi, Bernt SchieleICCV 2023 · 被引用 22 次
- Localizing Active Objects from Egocentric Vision with Symbolic World KnowledgeTe-Lin Wu, Yu Zhou, Nanyun PengEMNLP 2023 · 被引用 5 次
- AlignCAT: Visual-Linguistic Alignment of Category and Attribute for Weakly Supervised Visual GroundingYidan Wang, Chenyi Zhuang, Wutao Liu, Pan Gao 等ACM MM 2025 · 被引用 2 次
- Interpretable Composition Attribution Enhancement for Visio-linguistic Compositional UnderstandingWei Li, Zhen Huang, Xinmei Tian, Le Lu 等EMNLP 2024
它引用的顶会 Paper14
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty 等NeurIPS 2021 · 被引用 2,985 次
- Generic Attention-model Explainability for Interpreting Bi-Modal and Encoder-Decoder TransformersHila Chefer, Shir Gur, Lior WolfICCV 2021 · 被引用 451 次
- Taking a HINT: Leveraging Explanations to Make Vision and Language Models More GroundedRamprasaath Ramasamy Selvaraju, Stefan Lee, Yilin Shen, Hongxia Jin 等ICCV 2019 · 被引用 288 次
- Align2Ground: Weakly Supervised Phrase Grounding Guided by Image-Caption AlignmentSamyak Datta, Karan Sikka, Anirban Roy, Karuna Ahuja 等ICCV 2019 · 被引用 113 次
相关 Paper
- Improved Visual Grounding through Self-Consistent ExplanationsRuozhen He, Paola Cascante-Bonilla, Ziyan Yang, Alexander C. Berg 等CVPR 2024
- Contrastive Learning with Expectation-Maximization for Weakly Supervised Phrase GroundingKeqin Chen, Richong Zhang, Samuel Mensah, Yongyi MaoEMNLP 2022 · 被引用 3 次
- Mask Grounding for Referring Image SegmentationYong Xien Chng, Henry Zheng, Yizeng Han, Xuchong Qiu 等CVPR 2024
- Unveiling Parts Beyond Objects: Towards Finer-Granularity Referring Expression SegmentationWenxuan Wang, Tongtian Yue, Yisi Zhang, Longteng Guo 等CVPR 2024 · 被引用 5 次
- Multi-task Visual Grounding with Coarse-to-Fine Consistency ConstraintsMing Dai, Jian Li, Jiedong Zhuang, Xian Zhang 等AAAI 2025 · 被引用 23 次
