AutoAttend: Automated Attention Representation Search
Chaoyu Guan, Xin Wang, Wenwu Zhu
Abstract
Self-attention mechanisms have been widely adopted in many machine learning areas, including Natural Language Processing (NLP) and Graph Representation Learning (GRL), etc. However, existing works heavily rely on hand-crafted design to obtain customized attention mechanisms. In this paper, we automate Key, Query and Value representation design, which is one of the most important steps to obtain effective self-attentions. We propose an automated self-attention representation model, AutoAttend, which can automatically search powerful attention representations for downstream tasks leveraging Neural Architecture Search (NAS). In particular, we design a tailored search space for attention representation automation, which is flexible to produce effective attention representation designs. Based on the design prior obtained from attention representations in previous works, we further regularize our search space to reduce the space complexity without the loss of expressivity. Moreover, we propose a novel context-aware parameter sharing mechanism considering special characteristics of each sub-architecture to provide more accurate architecture estimations when conducting parameter sharing in our tailored search space. Experiments show the superiority of our proposed AutoAttend model over previous state-of-the-arts on eight text classification tasks in NLP and four node classification tasks in GRL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1e1f588a-e99b-4558-9715-4f508a97eb67Cited by top-tier papers9
- Graph Differentiable Architecture Search with Structure LearningYijian Qin, Xin Wang, Zeyang Zhang, Wenwu ZhuNeurIPS 2021 · 52 citations
- Multimodal Continual Graph Learning with Neural Architecture SearchJie Cai, Xin Wang, Chaoyu Guan, Yateng Tang et al.WWW 2022 · 50 citations
- Dynamic Heterogeneous Graph Attention Neural Architecture SearchZeyang Zhang, Ziwei Zhang, Xin Wang, Yijian Qin et al.AAAI 2023 · 44 citations
- Adaptive Disentangled Transformer for Sequential RecommendationYipeng Zhang, Xin Wang, Hong Chen, Wenwu ZhuKDD 2023 · 32 citations
- Large-Scale Graph Neural Architecture SearchChaoyu Guan, Xin Wang, Hong Chen, Ziwei Zhang et al.ICML 2022 · 27 citations
Builds on13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- Hyper-SAGNN: a self-attention based graph neural network for hypergraphsRuochi Zhang, Yuesong Zou, Jian MaICLR 2020 · 228 citations
- Bridging the Gap between Sample-based and One-shot Neural Architecture Search with BONASHan Shi, Renjie Pi, Hang Xu, Zhenguo Li et al.NeurIPS 2020 · 148 citations
- AtomNAS: Fine-Grained End-to-End Neural Architecture SearchJieru Mei, Yingwei Li, Xiaochen Lian, Xiaojie Jin et al.ICLR 2020 · 110 citations
Related papers
- TextNAS: A Neural Architecture Search Space Tailored for Text RepresentationYujing Wang, Yaming Yang, Yiren Chen, Jing Bai et al.AAAI 2020 · 66 citations
- PSP: Progressive Space Pruning for Efficient Graph Neural Architecture SearchGuanghui Zhu, Wenjie Wang, Zhuoer Xu, Feng Cheng et al.ICDE 2022 · 5 citations
- Auto Learning AttentionBenteng Ma, Jing Zhang, Yong Xia, Dacheng TaoNeurIPS 2020 · 28 citations
- K-armed Bandit based Multi-Modal Network Architecture Search for Visual Question AnsweringYiyi Zhou, Rongrong Ji, Xiaoshuai Sun, Gen Luo et al.ACM MM 2020 · 25 citations
- AutoGT: Automated Graph Transformer Architecture SearchZizhao Zhang, Xin Wang, Chaoyu Guan, Ziwei Zhang et al.ICLR 2023
