Hypergraph Attention Networks for Multimodal Learning
Eun-Sol Kim, Woo-Young Kang, Kyoung-Woon On, Yu-Jung Heo, Byoung-Tak Zhang
摘要
One of the fundamental problems that arise in multimodal learning tasks is the disparity of information levels between different modalities. To resolve this problem, we propose Hypergraph Attention Networks (HANs), which define a common semantic space among the modalities with symbolic graphs and extract a joint representation of the modalities based on a co-attention map constructed in the semantic space. HANs follow the process: constructing the common semantic space with symbolic graphs of each modality, matching the semantics between sub-structures of the symbolic graphs, constructing co-attention maps between the graphs in the semantic space, and integrating the multimodal inputs using the co-attention maps to get the final joint representation. From the qualitative analysis with two Visual Question and Answering datasets, we discover that 1) the alignment of the information levels between the modalities is important, and 2) the symbolic graphs are very powerful ways to represent the information of the low-level signals in alignment. Moreover, HANs dramatically improve the state-of-the-art accuracy on the GQA dataset from 54.6% to 61.88% only using the symbolic information in quantitatively.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Tag-assisted Multimodal Sentiment Analysis under Uncertain Missing ModalitiesJiandian Zeng, Tianyi Liu, Jiantao ZhouSIGIR 2022 · 被引用 84 次
- CAT-Walk: Inductive Hypergraph Learning via Set WalksAli Behrouz, Farnoosh Hashemi, Sadaf Sadeghian, Margo I. SeltzerNeurIPS 2023 · 被引用 21 次
- Copresheaf Topological Neural Networks: A Generalized Deep Learning FrameworkMustafa Hajij, Lennart Bastian, Sarah Osentoski, Hardik Kabaria 等NeurIPS 2025 · 被引用 15 次
- Adaptive Hypergraph Neural Network for Multi-Person Pose EstimationXixia Xu, Qi Zou, Xue LinAAAI 2022 · 被引用 14 次
- Hypergraph Neural Architecture SearchWei Lin, Xu Peng, Zhengtao Yu, Taisong JinAAAI 2024 · 被引用 3 次
它引用的顶会 Paper2
相关 Paper
- Hierarchical Graph Attention Network for Few-shot Visual-Semantic LearningChengxiang Yin, Kun Wu, Zhengping Che, Bo Jiang 等ICCV 2021 · 被引用 11 次
- VQA-GNN: Reasoning with Multimodal Knowledge via Graph Neural Networks for Visual Question AnsweringYanan Wang, Michihiro Yasunaga, Hongyu Ren, Shinya Wada 等ICCV 2023 · 被引用 42 次
- Reasoning with Heterogeneous Graph Alignment for Video Question AnsweringPin Jiang, Yahong HanAAAI 2020 · 被引用 214 次
- Object Attribute Matters in Visual Question AnsweringPeize Li, Qingyi Si, Peng Fu, Zheng Lin 等AAAI 2024 · 被引用 1 次
- Graph-Based Reasoning over Heterogeneous External Knowledge for Commonsense Question AnsweringShangwen Lv, Daya Guo, Jingjing Xu, Duyu Tang 等AAAI 2020 · 被引用 224 次
