Core-to-Global Reasoning for Compositional Visual Question Answering
Hao Zhou, Tingjin Luo, Zhangqi Jiang
Abstract
Compositional visual question answering (Compositional VQA) needs to provide an answer to a compositional question, which requires the model to have advanced capabilities of multi-modal semantic understanding and logical reasoning. However, current VQA models mainly concentrate on enriching the visual representations of images and neglect the redundancy in the enriched information to bring some negative impacts. To enhance the value and availability of semantic features, we propose a novel core-to-global reasoning (CTGR) model for compositional VQA. The model first extracts both global features and core features from image and question through a feature embedding module. Then, to enhance the value of semantic features, we propose an information filtering module to align visual features and text features at the core semantic level and to filter out the redundancy carried by image and question features at the global semantic level, which can further strengthen cross-modal correlations. Besides, we design a novel core-to-global reasoning mechanism for multimodal fusion, which integrates content features from core learning and context features from global features for accurate answer predictions. Finally, extensive experimental results on GQA, GQA-sub, VQA2.0 and Visual7W demonstrate the effectiveness and superiority of CTGR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 56201632-c8ef-4265-99c0-08ebcd948d73Cited by top-tier papers1
Ask how each one uses itBuilds on12
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
- Language-Conditioned Graph Networks for Relational ReasoningRonghang Hu, Anna Rohrbach, Trevor Darrell, Kate SaenkoICCV 2019 · 183 citations
- Overcoming Language Priors in VQA via Decomposed Linguistic RepresentationsChenchen Jing, Yuwei Wu, Xiaoxun Zhang, Yunde Jia et al.AAAI 2020 · 115 citations
- Adversarial VQA: A New Benchmark for Evaluating the Robustness of VQA ModelsLinjie Li, Jie Lei, Zhe Gan, Jingjing LiuICCV 2021 · 99 citations
- Compact Trilinear Interaction for Visual Question AnsweringTuong Do, Huy Tran, Thanh-Toan Do, Erman Tjiputra et al.ICCV 2019 · 63 citations
Related papers
- Focal and Composed Vision-semantic Modeling for Visual Question AnsweringYudong Han, Yangyang Guo, Jianhua Yin, Meng Liu et al.ACM MM 2021 · 14 citations
- Maintaining Reasoning Consistency in Compositional Visual Question AnsweringChenchen Jing, Yunde Jia, Yuwei Wu, Xinyu Liu et al.CVPR 2022 · 27 citations
- Variational Causal Inference Network for Explanatory Visual Question AnsweringDizhan Xue, Shengsheng Qian, Changsheng XuICCV 2023 · 19 citations
- VQA-GNN: Reasoning with Multimodal Knowledge via Graph Neural Networks for Visual Question AnsweringYanan Wang, Michihiro Yasunaga, Hongyu Ren, Shinya Wada et al.ICCV 2023 · 42 citations
- Beyond OCR + VQA: Involving OCR into the Flow for Robust and Accurate TextVQAGangyan Zeng, Yuan Zhang, Yu Zhou, Xiaomeng YangACM MM 2021 · 38 citations
