Compact Trilinear Interaction for Visual Question Answering
Tuong Do, Huy Tran, Thanh-Toan Do, Erman Tjiputra, Quang D. Tran
摘要
In Visual Question Answering (VQA), answers have a great correlation with question meaning and visual contents. Thus, to selectively utilize image, question and answer information, we propose a novel trilinear interaction model which simultaneously learns high level associations between these three inputs. In addition, to overcome the interaction complexity, we introduce a multimodal tensor-based PARALIND decomposition which efficiently parameterizes trilinear interaction between the three inputs. Moreover, knowledge distillation is first time applied in Free-form Opened-ended VQA. It is not only for reducing the computational cost and required memory but also for transferring knowledge from trilinear interaction model to bilinear interaction model. The extensive experiments on benchmarking datasets TDIUC, VQA-2.0, and Visual7W show that the proposed compact trilinear interaction model achieves state-of-the-art results when using a single model on all three datasets. The source code is available at https://github.com/aioz-ai/ICCV19_ VQA-CTI .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- VidLanKD: Improving Language Understanding via Video-Distilled Knowledge TransferZineng Tang, Jaemin Cho, Hao Tan, Mohit BansalNeurIPS 2021 · 被引用 36 次
- PEVL: Position-enhanced Pre-training and Prompt Tuning for Vision-language ModelsYuan Yao, Qianyu Chen, Ao Zhang, Wei Ji 等EMNLP 2022 · 被引用 33 次
- Unsupervised Cross-Modal Distillation for Thermal Infrared TrackingJingxian Sun, Lichao Zhang, Yufei Zha, Abel Gonzalez-Garcia 等ACM MM 2021 · 被引用 32 次
- Depth Privileged Object Detection in Indoor Scenes via Deformation HallucinationZhijie Zhang, Yan Liu, Junjie Chen, Li Niu 等AAAI 2021 · 被引用 6 次
- Core-to-Global Reasoning for Compositional Visual Question AnsweringHao Zhou, Tingjin Luo, Zhangqi JiangAAAI 2025 · 被引用 2 次
相关 Paper
- Object Attribute Matters in Visual Question AnsweringPeize Li, Qingyi Si, Peng Fu, Zheng Lin 等AAAI 2024 · 被引用 1 次
- Multi-Modality Latent Interaction Network for Visual Question AnsweringPeng Gao, Haoxuan You, Zhanpeng Zhang, Xiaogang Wang 等ICCV 2019 · 被引用 86 次
- Decouple Before Interact: Multi-Modal Prompt Learning for Continual Visual Question AnsweringZi Qian, Xin Wang, Xuguang Duan, Pengda Qin 等ICCV 2023 · 被引用 28 次
- From Superficial to Deep: Language Bias driven Curriculum Learning for Visual Question AnsweringMingrui Lao, Yanming Guo, Yu Liu, Wei Chen 等ACM MM 2021 · 被引用 22 次
- KAID: Knowledge-Aware Interactive Distillation for Vision-Language ModelsDa Zhang, Feiyu Wang, Bingyu Li, Zhiyuan Zhao 等ACM MM 2025 · 被引用 10 次
