Grounded Image Text Matching with Mismatched Relation Reasoning
Yu Wu, Yana Wei, Haozhe Wang, Yongfei Liu, Sibei Yang, Xuming He
摘要
This paper introduces Grounded Image Text Matching with Mismatched Relation (GITM-MR), a novel visual-linguistic joint task that evaluates the relation understanding capabilities of transformer-based pre-trained models. GITM-MR requires a model to first determine if an expression describes an image, then localize referred objects or ground the mismatched parts of the text. We provide a benchmark for evaluating vision-language (VL) models on this task, with a focus on the challenging settings of limited training data and out-of-distribution sentence lengths. Our evaluation demonstrates that pre-trained VL models often lack data efficiency and length generalization ability. To address this, we propose the Relation-sensitive Correspondence Reasoning Network (RCRN), which incorporates relation-aware reasoning via bi-directional message propagation guided by language structure. Our RCRN can he interpreted as a modular program and delivers strong performance in terms of both length generalization and data efficiency. The code and data are available on https://githuh.coin/SHTUPLUS/GITM-MR.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Viewpoint-Aware Visual Grounding in 3D ScenesXiangxi Shi, Zhonghua Wu, Stefan LeeCVPR 2024 · 被引用 13 次
- Intervene-All-Paths: Unified Mitigation of LVLM Hallucinations across Alignment FormatsJiaye Qian, Ge Zheng, Yuchen Zhu, Sibei YangNeurIPS 2025 · 被引用 11 次
- Why LVLMs are More Prone to Hallucinations in Longer Responses: The Role of ContextGe Zheng, Jiaye Qian, Jiajin Tang, Sibei YangICCV 2025 · 被引用 2 次
- FineCops-Ref: A new Dataset and Task for Fine-Grained Compositional Referring Expression ComprehensionJunzhuo Liu, Xuzheng Yang, Weiwei Li, Peng WangEMNLP 2024 · 被引用 2 次
- Gradient-Attention Guided Dual-Masking Synergetic Framework for Robust Text-based Person RetrievalTianlu Zheng, Yifan Zhang, Xiang An, Ziyong Feng 等EMNLP 2025 · 被引用 1 次
它引用的顶会 Paper30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 被引用 1,168 次
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve 等ICCV 2021 · 被引用 1,114 次
- Visual Semantic Reasoning for Image-Text MatchingKunpeng Li, Yulun Zhang, Kai Li, Yuanyuan Li 等ICCV 2019 · 被引用 598 次
- Large-Scale Adversarial Training for Vision-and-Language Representation LearningZhe Gan, Yen-Chun Chen, Linjie Li, Chen Zhu 等NeurIPS 2020 · 被引用 561 次
相关 Paper
- Do Vision-Language Transformers Exhibit Visual Commonsense? An Empirical Study of VCRZhenyang Li, Yangyang Guo, Kejie Wang, Xiaolin Chen 等ACM MM 2023 · 被引用 11 次
- Advancing Visual Grounding with Scene Knowledge: Benchmark and MethodZhihong Chen, Ruifei Zhang, Yibing Song, Xiang Wan 等CVPR 2023
- Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-TrainingGen Li, Nan Duan, Yuejian Fang, Ming Gong 等AAAI 2020 · 被引用 966 次
- RelViT: Concept-guided Vision Transformer for Visual Relational ReasoningXiaojian Ma, Weili Nie, Zhiding Yu, Huaizu Jiang 等ICLR 2022 · 被引用 21 次
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li 等ICLR 2020 · 被引用 1,825 次
