Multi-Modality Latent Interaction Network for Visual Question Answering
Peng Gao, Haoxuan You, Zhanpeng Zhang, Xiaogang Wang, Hongsheng Li
摘要
Exploiting relationships between visual regions and question words have achieved great success in learning multi-modality features for Visual Question Answering (VQA). However, we argue that existing methods [29] mostly model relations between individual visual regions and words, which are not enough to correctly answer the question. From humans' perspective, answering a visual question requires understanding the summarizations of visual and language information. In this paper, we proposed the Multi-modality Latent Interaction module (MLI) to tackle this problem. The proposed module learns the cross-modality relationships between latent visual and language summarizations, which summarize visual regions and question into a small number of latent representations to avoid modeling uninformative individual region-word relations. The cross-modality information between the latent summarizations are propagated to fuse valuable information from both modalities and are used to update the visual and word features. Such MLI modules can be stacked for several stages to model complex and latent relations between the two modalities and achieves highly competitive performance on public VQA benchmarks, VQA v2.0 [12] and TDIUC [20] . In addition, we show that the performance of our methods could be significantly improved by combining with pre-trained language model BERT[6].
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Greedy Gradient Ensemble for Robust Visual Question AnsweringXinzhe Han, Shuhui Wang, Chi Su, Qingming Huang 等ICCV 2021 · 被引用 94 次
- Dual-stream Network for Visual RecognitionMingyuan Mao, Peng Gao, Renrui Zhang, Honghui Zheng 等NeurIPS 2021 · 被引用 88 次
- Container: Context Aggregation NetworksPeng Gao, Jiasen Lu, Hongsheng Li, Roozbeh Mottaghi 等NeurIPS 2021 · 被引用 86 次
- Unshuffling Data for Improved Generalization in Visual Question AnsweringDamien Teney, Ehsan Abbasnejad, Anton van den HengelICCV 2021 · 被引用 84 次
- Fastened CROWN: Tightened Neural Network Robustness CertificatesZhaoyang Lyu, Ching-Yun Ko, Zhifeng Kong, Ngai Wong 等AAAI 2020 · 被引用 70 次
相关 Paper
- Cross-modal Information Flow in Multimodal Large Language ModelsZhi Zhang, Srishti Yadav, Fengze Han, Ekaterina ShutovaCVPR 2025
- Cross-Modality Relevance for Reasoning on Language and VisionChen Zheng, Quan Guo, Parisa KordjamshidiACL 2020 · 被引用 33 次
- Pairwise VLAD Interaction Network for Video Question AnsweringHui Wang, Dan Guo, Xian-Sheng Hua, Meng WangACM MM 2021 · 被引用 15 次
- Modality Eigen-Encodings Are Keys to Open Modality Informative ContainersYiyuan Zhang, Yuqi JiACM MM 2022
- Multi-modal Data Spectrum: Multi-modal Datasets are Multi-dimensionalDivyam Madaan, Varshan Muhunthan, Kyunghyun Cho, Sumit ChopraICLR 2026 · 被引用 3 次
