Query and Attention Augmentation for Knowledge-Based Explainable Reasoning
Yifeng Zhang, Ming Jiang, Qi Zhao
摘要
Explainable visual question answering (VQA) models have been developed with neural modules and query-based knowledge incorporation to answer knowledge-requiring questions. Yet, most reasoning methods cannot effectively generate queries or incorporate external knowledge during the reasoning process, which may lead to suboptimal results. To bridge this research gap, we present Query and Attention Augmentation, a general approach that augments neural module networks to jointly reason about visual and external knowledge. To take both knowledge sources into account during reasoning, it parses the input question into a functional program with queries augmented through a novel reinforcement learning method, and jointly directs augmented attention to visual and external knowledge based on intermediate reasoning results. With extensive experiments on multiple VQA datasets, our method demonstrates significant performance, explainability, and generalizability over state-of-the-art models in answering questions requiring different extents of knowledge. Our source code is available at https://github.com/SuperJohnZhang/QAA .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Overcoming Dual Drift for Continual Long-Tailed Visual Question AnsweringFeifei Zhang, Zhihao Wang, Xi Zhang, Changsheng XuICCV 2025 · 被引用 3 次
- PAE: Reinforcement Learning from External Knowledge for Efficient ExplorationZhe Wu, Haofei Lu, Junliang Xing, You Wu 等ICLR 2024 · 被引用 1 次
- VQACL: A Novel Visual Question Answering Continual Learning SettingXi Zhang, Feifei Zhang, Changsheng XuCVPR 2023
它引用的顶会 Paper3
- Learning to Collocate Neural Modules for Image CaptioningXu Yang, Hanwang Zhang, Jianfei CaiICCV 2019 · 被引用 84 次
- Explicit Knowledge Incorporation for Visual ReasoningYifeng Zhang, Ming Jiang, Qi ZhaoCVPR 2021
- Fantastic Answers and Where to Find Them: Immersive Question-Directed Visual AttentionMing Jiang, Shi Chen, Jinhui Yang, Qi ZhaoCVPR 2020
相关 Paper
- ReAG: Reasoning-Augmented Generation for Knowledge-based Visual Question AnsweringAlberto Compagnoni, Marco Morini, Sara Sarto, Federico Cocchi 等CVPR 2026 · 被引用 11 次
- Retrieval-Augmented Visual Question Answering via Built-in Autoregressive Search EnginesXinwei Long, Zhiyuan Ma, Ermo Hua, Kaiyan Zhang 等AAAI 2025 · 被引用 18 次
- Boosting Visual Question Answering with Context-aware Knowledge AggregationGuohao Li, Xin Wang, Wenwu ZhuACM MM 2020 · 被引用 82 次
- MMRAG-RFT: Two-stage Reinforcement Fine-tuning for Explainable Multi-modal Retrieval-augmented GenerationShengwei Zhao, Jingwen Yao, Sitong Wei, Linhai Xu 等AAAI 2026
- Large Language Models Know What is Key Visual Entity: An LLM-assisted Multimodal Retrieval for VQAPu Jian, Donglei Yu, Jiajun ZhangEMNLP 2024 · 被引用 5 次
