Multi-Modality Latent Interaction Network for Visual Question Answering
Peng Gao, Haoxuan You, Zhanpeng Zhang, Xiaogang Wang, Hongsheng Li
Abstract
Exploiting relationships between visual regions and question words have achieved great success in learning multi-modality features for Visual Question Answering (VQA). However, we argue that existing methods [29] mostly model relations between individual visual regions and words, which are not enough to correctly answer the question. From humans' perspective, answering a visual question requires understanding the summarizations of visual and language information. In this paper, we proposed the Multi-modality Latent Interaction module (MLI) to tackle this problem. The proposed module learns the cross-modality relationships between latent visual and language summarizations, which summarize visual regions and question into a small number of latent representations to avoid modeling uninformative individual region-word relations. The cross-modality information between the latent summarizations are propagated to fuse valuable information from both modalities and are used to update the visual and word features. Such MLI modules can be stacked for several stages to model complex and latent relations between the two modalities and achieves highly competitive performance on public VQA benchmarks, VQA v2.0 [12] and TDIUC [20] . In addition, we show that the performance of our methods could be significantly improved by combining with pre-trained language model BERT[6].
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8b72b310-4b7c-4b7f-8a1b-36fdfb9ebdfaCited by top-tier papers11
- Greedy Gradient Ensemble for Robust Visual Question AnsweringXinzhe Han, Shuhui Wang, Chi Su, Qingming Huang et al.ICCV 2021 · 94 citations
- Dual-stream Network for Visual RecognitionMingyuan Mao, Peng Gao, Renrui Zhang, Honghui Zheng et al.NeurIPS 2021 · 88 citations
- Container: Context Aggregation NetworksPeng Gao, Jiasen Lu, Hongsheng Li, Roozbeh Mottaghi et al.NeurIPS 2021 · 86 citations
- Unshuffling Data for Improved Generalization in Visual Question AnsweringDamien Teney, Ehsan Abbasnejad, Anton van den HengelICCV 2021 · 84 citations
- Fastened CROWN: Tightened Neural Network Robustness CertificatesZhaoyang Lyu, Ching-Yun Ko, Zhifeng Kong, Ngai Wong et al.AAAI 2020 · 70 citations
Related papers
- Cross-modal Information Flow in Multimodal Large Language ModelsZhi Zhang, Srishti Yadav, Fengze Han, Ekaterina ShutovaCVPR 2025
- Cross-Modality Relevance for Reasoning on Language and VisionChen Zheng, Quan Guo, Parisa KordjamshidiACL 2020 · 33 citations
- Pairwise VLAD Interaction Network for Video Question AnsweringHui Wang, Dan Guo, Xian-Sheng Hua, Meng WangACM MM 2021 · 15 citations
- Modality Eigen-Encodings Are Keys to Open Modality Informative ContainersYiyuan Zhang, Yuqi JiACM MM 2022
- Multi-modal Data Spectrum: Multi-modal Datasets are Multi-dimensionalDivyam Madaan, Varshan Muhunthan, Kyunghyun Cho, Sumit ChopraICLR 2026 · 3 citations
