Calibrating Concepts and Operations: Towards Symbolic Reasoning on Real Images
Zhuowan Li, Elias Stengel-Eskin, Yixiao Zhang, Cihang Xie, Quan Tran, Benjamin Van Durme, Alan L. Yuille
摘要
While neural symbolic methods demonstrate impressive performance in visual question answering on synthetic images, their performance suffers on real images. We identify that the long-tail distribution of visual concepts and unequal importance of reasoning steps in real data are the two key obstacles that limit the models’ real-world potentials. To address these challenges, we propose a new paradigm, Calibrating Concepts and Operations (CCO), which enables neural symbolic models to capture underlying data characteristics and to reason with hierarchical importance. Specifically, we introduce an executor with learnable concept embedding magnitudes for handling distribution imbalance, and an operation calibrator for highlighting important operations and suppressing redundant ones.Our experiments show CCO substantially boosts the performance of neural symbolic methods on real images. By evaluating models on the real world dataset GQA, CCO helps the neural symbolic method NSCL outperforms its vanilla counterpart by 9.1% (from 47.0% to 56.1%); this result also largely reduces the performance gap between symbolic and non-symbolic methods. Additionally, we create a perturbed test set for better understanding and analyzing model performance on real images. Code is available at https://lizw14.github.io/project/ccosr.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Look, Remember and Reason: Grounded Reasoning in Videos with Language ModelsApratim Bhattacharyya, Sunny Panchal, Reza Pourreza, Mingu Lee 等ICLR 2024 · 被引用 15 次
- Super-CLEVR: A Virtual Benchmark to Diagnose Domain Robustness in Visual ReasoningZhuowan Li, Xingrui Wang, Elias Stengel-Eskin, Adam Kortylewski 等CVPR 2023
- Synthesize Step-by-Step: Tools, Templates and LLMs as Data Generators for Reasoning-Based Chart VQAZhuowan Li, Bhavan Jasani, Peng Tang, Shabnam GhadarCVPR 2024
它引用的顶会 Paper8
- Dynamic Graph Attention for Referring Expression ComprehensionSibei Yang, Guanbin Li, Yizhou YuICCV 2019 · 被引用 251 次
- Neuro-Symbolic Visual Reasoning: Disentangling "Visual" from "Reasoning"Saeed Amizadeh, Hamid Palangi, Alex Polozov, Yichen Huang 等ICML 2020 · 被引用 74 次
- Machine Number Sense: A Dataset of Visual Arithmetic Problems for Abstract and Relational ReasoningWenhe Zhang, Chi Zhang, Yixin Zhu, Song-Chun ZhuAAAI 2020 · 被引用 31 次
- Counterfactual Samples Synthesizing for Robust Visual Question AnsweringLong Chen, Xin Yan, Jun Xiao, Hanwang Zhang 等CVPR 2020
- ACRE: Abstract Causal REasoning Beyond CovariationChi Zhang, Baoxiong Jia, Mark Edmonds, Song-Chun Zhu 等CVPR 2021
相关 Paper
- Interpretable Visual Reasoning via Induced Symbolic SpaceZhonghao Wang, Kai Wang, Mo Yu, Jinjun Xiong 等ICCV 2021 · 被引用 22 次
- NeuS-QA: Grounding Long-Form Video Understanding in Temporal Logic and Neuro-Symbolic ReasoningSahil Shah, S. P. Sharan, Harsh Goel, Minkyu Choi 等AAAI 2026 · 被引用 4 次
- Image Manipulation via Multi-Hop Instructions - A New Dataset and Weakly-Supervised Neuro-Symbolic ApproachHarman Singh, Poorva Garg, Mohit Gupta, Kevin Shah 等EMNLP 2023
- Naturally Supervised 3D Visual Grounding with Language-Regularized Concept LearnersChun Feng, Joy Hsu, Weiyu Liu, Jiajun WuCVPR 2024
- V-PROM: A Benchmark for Visual Reasoning Using Visual Progressive MatricesDamien Teney, Peng Wang, Jiewei Cao, Lingqiao Liu 等AAAI 2020 · 被引用 37 次
