K-armed Bandit based Multi-Modal Network Architecture Search for Visual Question Answering
Yiyi Zhou, Rongrong Ji, Xiaoshuai Sun, Gen Luo, Xiaopeng Hong, Jinsong Su, Xinghao Ding, Ling Shao
Abstract
In this paper, we propose a cross-modal network architecture search (NAS) algorithm for VQA, termed as k-Armed Bandit based NAS (KAB-NAS). KAB-NAS regards the design of each layer as a k-armed bandit problem and updates the preference of each candidate via numerous samplings in a single-shot search framework. To establish an effective search space, we further propose a new architecture termed Automatic Graph Attention Network (AGAN), and extend the popular self-attention layer with three graph structures, denoted as dense-graph, co-graph and separate-graph.These graph layers are used to form the direction of information propagation in the graph network, and their optimal combinations are searched by KAB-NAS. To evaluate KAB-NAS and AGAN, we conduct extensive experiments on two VQA benchmark datasets, i.e., VQA2.0 and GQA, and also test AGAN with the popular BERT-style pre-training. The experimental results show that with the help of KAB-NAS, AGAN can achieve the state-of-the-art performance on both benchmark datasets with much fewer parameters and computations.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers4
- TRAR: Routing the Attention Spans in Transformer for Visual Question AnsweringYiyi Zhou, Tianhe Ren, Chaoyang Zhu, Xiaoshuai Sun et al.ICCV 2021 · 128 citations
- Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language ModelsWeihao Ye, Qiong Wu, Wenhao Lin, Yiyi ZhouAAAI 2025 · 99 citations
- Exploiting the Social-Like Prior in Transformer for Visual ReasoningYudong Han, Yupeng Hu, Xuemeng Song, Haoyu Tang et al.AAAI 2024 · 11 citations
- Combinatorial Pure Exploration with Bottleneck Reward FunctionYihan Du, Yuko Kuroki, Wei ChenNeurIPS 2021 · 6 citations
Related papers
- AGNAS: Attention-Guided Micro and Macro-Architecture SearchZihao Sun, Yu Hu, Shun Lu, Longxing Yang et al.ICML 2022 · 16 citations
- AutoAttend: Automated Attention Representation SearchChaoyu Guan, Xin Wang, Wenwu ZhuICML 2021 · 46 citations
- Customizing Graph Neural Network for CAD Assembly RecommendationFengqi Liang, Huan Zhao, Yuhan Quan, Wei Fang et al.KDD 2024 · 3 citations
- Deep Multimodal Neural Architecture SearchZhou Yu, Yuhao Cui, Jun Yu, Meng Wang et al.ACM MM 2020 · 93 citations
- Not All Operations Contribute Equally: Hierarchical Operation-adaptive Predictor for Neural Architecture SearchZiye Chen, Yibing Zhan, Baosheng Yu, Mingming Gong et al.ICCV 2021 · 13 citations
