Compact Trilinear Interaction for Visual Question Answering
Tuong Do, Huy Tran, Thanh-Toan Do, Erman Tjiputra, Quang D. Tran
Abstract
In Visual Question Answering (VQA), answers have a great correlation with question meaning and visual contents. Thus, to selectively utilize image, question and answer information, we propose a novel trilinear interaction model which simultaneously learns high level associations between these three inputs. In addition, to overcome the interaction complexity, we introduce a multimodal tensor-based PARALIND decomposition which efficiently parameterizes trilinear interaction between the three inputs. Moreover, knowledge distillation is first time applied in Free-form Opened-ended VQA. It is not only for reducing the computational cost and required memory but also for transferring knowledge from trilinear interaction model to bilinear interaction model. The extensive experiments on benchmarking datasets TDIUC, VQA-2.0, and Visual7W show that the proposed compact trilinear interaction model achieves state-of-the-art results when using a single model on all three datasets. The source code is available at https://github.com/aioz-ai/ICCV19_ VQA-CTI .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 70237dce-2f9e-4b83-800f-5db232882783Cited by top-tier papers9
- VidLanKD: Improving Language Understanding via Video-Distilled Knowledge TransferZineng Tang, Jaemin Cho, Hao Tan, Mohit BansalNeurIPS 2021 · 36 citations
- PEVL: Position-enhanced Pre-training and Prompt Tuning for Vision-language ModelsYuan Yao, Qianyu Chen, Ao Zhang, Wei Ji et al.EMNLP 2022 · 33 citations
- Unsupervised Cross-Modal Distillation for Thermal Infrared TrackingJingxian Sun, Lichao Zhang, Yufei Zha, Abel Gonzalez-Garcia et al.ACM MM 2021 · 32 citations
- Depth Privileged Object Detection in Indoor Scenes via Deformation HallucinationZhijie Zhang, Yan Liu, Junjie Chen, Li Niu et al.AAAI 2021 · 6 citations
- Core-to-Global Reasoning for Compositional Visual Question AnsweringHao Zhou, Tingjin Luo, Zhangqi JiangAAAI 2025 · 2 citations
Related papers
- Object Attribute Matters in Visual Question AnsweringPeize Li, Qingyi Si, Peng Fu, Zheng Lin et al.AAAI 2024 · 1 citation
- Multi-Modality Latent Interaction Network for Visual Question AnsweringPeng Gao, Haoxuan You, Zhanpeng Zhang, Xiaogang Wang et al.ICCV 2019 · 86 citations
- Decouple Before Interact: Multi-Modal Prompt Learning for Continual Visual Question AnsweringZi Qian, Xin Wang, Xuguang Duan, Pengda Qin et al.ICCV 2023 · 28 citations
- From Superficial to Deep: Language Bias driven Curriculum Learning for Visual Question AnsweringMingrui Lao, Yanming Guo, Yu Liu, Wei Chen et al.ACM MM 2021 · 22 citations
- KAID: Knowledge-Aware Interactive Distillation for Vision-Language ModelsDa Zhang, Feiyu Wang, Bingyu Li, Zhiyuan Zhao et al.ACM MM 2025 · 10 citations
