Interpretable Visual Reasoning via Induced Symbolic Space
Zhonghao Wang, Kai Wang, Mo Yu, Jinjun Xiong, Wen-Mei Hwu, Mark Hasegawa-Johnson, Humphrey Shi
摘要
We study the problem of concept induction in visual reasoning, i.e., identifying concepts and their hierarchical relationships from question-answer pairs associated with images; and achieve an interpretable model via working on the induced symbolic concept space. To this end, we first design a new framework named object-centric compositional attention model (OCCAM) to perform the visual reasoning task with object-level visual features. Then, we come up with a method to induce concepts of objects and relations using clues from the attention patterns between objects’ visual features and question words. Finally, we achieve a higher level of interpretability by imposing OCCAM on the objects represented in the induced symbolic concept space. Experiments on the CLEVR and GQA datasets demonstrate: 1) our OCCAM achieves a new state of the art without human-annotated functional programs; 2) our induced concepts are both accurate and sufficient as OCCAM achieves an on-par performance on objects represented either in visual features or in the induced symbolic concept space.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve 等ICCV 2021 · 被引用 1,114 次
- CrossPoint: Self-Supervised Cross-Modal Contrastive Learning for 3D Point Cloud UnderstandingMohamed Afham, Isuru Dissanayake, Dinithi Dissanayake, Amaya Dharmasiri 等CVPR 2022 · 被引用 286 次
- Versatile Diffusion: Text, Images and Variations All in One Diffusion ModelXingqian Xu, Zhangyang Wang, Eric J. Zhang, Kai Wang 等ICCV 2023 · 被引用 265 次
- Exploiting the Social-Like Prior in Transformer for Visual ReasoningYudong Han, Yupeng Hu, Xuemeng Song, Haoyu Tang 等AAAI 2024 · 被引用 11 次
它引用的顶会 Paper3
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 被引用 2,258 次
- Language-Conditioned Graph Networks for Relational ReasoningRonghang Hu, Anna Rohrbach, Trevor Darrell, Kate SaenkoICCV 2019 · 被引用 183 次
- G3raphGround: Graph-Based Language GroundingMohit Bajaj, Lanjun Wang, Leonid SigalICCV 2019 · 被引用 67 次
相关 Paper
- Systematic Visual Reasoning through Object-Centric Relational AbstractionTaylor W. Webb, Shanka Subhra Mondal, Jonathan D. CohenNeurIPS 2023 · 被引用 35 次
- Neuro-Symbolic Visual Reasoning: Disentangling "Visual" from "Reasoning"Saeed Amizadeh, Hamid Palangi, Alex Polozov, Yichen Huang 等ICML 2020 · 被引用 74 次
- Decomposition of Concept-Level Rules in Visual ScenesFan Shi, Yuxuan Liang, Xiaolei Chen, Haiyang Yu 等ICLR 2026
- Interpretable Composition Attribution Enhancement for Visio-linguistic Compositional UnderstandingWei Li, Zhen Huang, Xinmei Tian, Le Lu 等EMNLP 2024
- Maintaining Reasoning Consistency in Compositional Visual Question AnsweringChenchen Jing, Yunde Jia, Yuwei Wu, Xinyu Liu 等CVPR 2022 · 被引用 27 次
