RelViT: Concept-guided Vision Transformer for Visual Relational Reasoning
Xiaojian Ma, Weili Nie, Zhiding Yu, Huaizu Jiang, Chaowei Xiao, Yuke Zhu, Song-Chun Zhu, Anima Anandkumar
摘要
Reasoning about visual relationships is central to how humans interpret the visual world. This task remains challenging for current deep learning algorithms since it requires addressing three key technical problems jointly: 1) identifying object entities and their properties, 2) inferring semantic relations between pairs of entities, and 3) generalizing to novel object-relation combinations, i.e. systematic generalization. In this work, we use vision transformers (ViTs) as our base model for visual reasoning and make better use of concepts defined as object entities and their relations to improve the reasoning ability of ViTs. Specifically, we introduce a novel concept-feature dictionary to allow flexible image feature retrieval at training time with concept keys. This dictionary enables two new conceptguided auxiliary tasks: 1) a global task for promoting relational reasoning, and 2) a local task for facilitating semantic object-centric correspondence learning. To examine the systematic generalization of visual reasoning models, we introduce systematic splits for the standard HICO and GQA benchmarks. We show the resulting model, Concept-guided Vision Transformer (or RelViT for short) significantly outperforms prior approaches on HICO and GQA by 16% and 13% in the original split, and by 43% and 18% in the systematic split. Our ablation analyses also reveal our model's compatibility with multiple ViT variants and robustness to hyper-parameters. Code is available.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Symbol-LLM: Leverage Language Models for Symbolic System in Visual Human Activity ReasoningXiaoqian Wu, Yonglu Li, Jianhua Sun, Cewu LuNeurIPS 2023 · 被引用 40 次
- Bongard-HOI: Benchmarking Few-Shot Visual Reasoning for Human-Object InteractionsHuaizu Jiang, Xiaojian Ma, Weili Nie, Zhiding Yu 等CVPR 2022 · 被引用 22 次
- Bongard-OpenWorld: Few-Shot Reasoning for Free-form Visual Concepts in the Real WorldRujie Wu, Xiaojian Ma, Zhenliang Zhang, Wei Wang 等ICLR 2024 · 被引用 20 次
- SQA3D: Situated Question Answering in 3D ScenesXiaojian Ma, Silong Yong, Zilong Zheng, Qing Li 等ICLR 2023 · 被引用 16 次
- Open-Set Image Tagging with Multi-Grained Text SupervisionXinyu Huang, Yi-Jie Huang, Youcai Zhang, Weiwei Tian 等ACM MM 2025 · 被引用 13 次
它引用的顶会 Paper22
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan 等ICCV 2021 · 被引用 4,909 次
相关 Paper
- Beyond the Doors of Perception: Vision Transformers Represent Relations Between ObjectsMichael A. Lepori, Alexa R. Tartaglini, Wai Keen Vong, Thomas Serre 等NeurIPS 2024 · 被引用 22 次
- Do Vision-Language Transformers Exhibit Visual Commonsense? An Empirical Study of VCRZhenyang Li, Yangyang Guo, Kejie Wang, Xiaolin Chen 等ACM MM 2023 · 被引用 11 次
- Slot Abstractors: Toward Scalable Abstract Visual ReasoningShanka Subhra Mondal, Jonathan D. Cohen, Taylor Whittington WebbICML 2024 · 被引用 10 次
- Weakly Supervised Relative Spatial Reasoning for Visual Question AnsweringPratyay Banerjee, Tejas Gokhale, Yezhou Yang, Chitta BaralICCV 2021 · 被引用 19 次
- Grounded Image Text Matching with Mismatched Relation ReasoningYu Wu, Yana Wei, Haozhe Wang, Yongfei Liu 等ICCV 2023 · 被引用 14 次
