K-armed Bandit based Multi-Modal Network Architecture Search for Visual Question Answering
Yiyi Zhou, Rongrong Ji, Xiaoshuai Sun, Gen Luo, Xiaopeng Hong, Jinsong Su, Xinghao Ding, Ling Shao
摘要
In this paper, we propose a cross-modal network architecture search (NAS) algorithm for VQA, termed as k-Armed Bandit based NAS (KAB-NAS). KAB-NAS regards the design of each layer as a k-armed bandit problem and updates the preference of each candidate via numerous samplings in a single-shot search framework. To establish an effective search space, we further propose a new architecture termed Automatic Graph Attention Network (AGAN), and extend the popular self-attention layer with three graph structures, denoted as dense-graph, co-graph and separate-graph.These graph layers are used to form the direction of information propagation in the graph network, and their optimal combinations are searched by KAB-NAS. To evaluate KAB-NAS and AGAN, we conduct extensive experiments on two VQA benchmark datasets, i.e., VQA2.0 and GQA, and also test AGAN with the popular BERT-style pre-training. The experimental results show that with the help of KAB-NAS, AGAN can achieve the state-of-the-art performance on both benchmark datasets with much fewer parameters and computations.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper4
- TRAR: Routing the Attention Spans in Transformer for Visual Question AnsweringYiyi Zhou, Tianhe Ren, Chaoyang Zhu, Xiaoshuai Sun 等ICCV 2021 · 被引用 128 次
- Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language ModelsWeihao Ye, Qiong Wu, Wenhao Lin, Yiyi ZhouAAAI 2025 · 被引用 99 次
- Exploiting the Social-Like Prior in Transformer for Visual ReasoningYudong Han, Yupeng Hu, Xuemeng Song, Haoyu Tang 等AAAI 2024 · 被引用 11 次
- Combinatorial Pure Exploration with Bottleneck Reward FunctionYihan Du, Yuko Kuroki, Wei ChenNeurIPS 2021 · 被引用 6 次
相关 Paper
- AGNAS: Attention-Guided Micro and Macro-Architecture SearchZihao Sun, Yu Hu, Shun Lu, Longxing Yang 等ICML 2022 · 被引用 16 次
- AutoAttend: Automated Attention Representation SearchChaoyu Guan, Xin Wang, Wenwu ZhuICML 2021 · 被引用 46 次
- Customizing Graph Neural Network for CAD Assembly RecommendationFengqi Liang, Huan Zhao, Yuhan Quan, Wei Fang 等KDD 2024 · 被引用 3 次
- Deep Multimodal Neural Architecture SearchZhou Yu, Yuhao Cui, Jun Yu, Meng Wang 等ACM MM 2020 · 被引用 93 次
- Not All Operations Contribute Equally: Hierarchical Operation-adaptive Predictor for Neural Architecture SearchZiye Chen, Yibing Zhan, Baosheng Yu, Mingming Gong 等ICCV 2021 · 被引用 13 次
