AutoAttend: Automated Attention Representation Search
Chaoyu Guan, Xin Wang, Wenwu Zhu
摘要
Self-attention mechanisms have been widely adopted in many machine learning areas, including Natural Language Processing (NLP) and Graph Representation Learning (GRL), etc. However, existing works heavily rely on hand-crafted design to obtain customized attention mechanisms. In this paper, we automate Key, Query and Value representation design, which is one of the most important steps to obtain effective self-attentions. We propose an automated self-attention representation model, AutoAttend, which can automatically search powerful attention representations for downstream tasks leveraging Neural Architecture Search (NAS). In particular, we design a tailored search space for attention representation automation, which is flexible to produce effective attention representation designs. Based on the design prior obtained from attention representations in previous works, we further regularize our search space to reduce the space complexity without the loss of expressivity. Moreover, we propose a novel context-aware parameter sharing mechanism considering special characteristics of each sub-architecture to provide more accurate architecture estimations when conducting parameter sharing in our tailored search space. Experiments show the superiority of our proposed AutoAttend model over previous state-of-the-arts on eight text classification tasks in NLP and four node classification tasks in GRL.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Graph Differentiable Architecture Search with Structure LearningYijian Qin, Xin Wang, Zeyang Zhang, Wenwu ZhuNeurIPS 2021 · 被引用 52 次
- Multimodal Continual Graph Learning with Neural Architecture SearchJie Cai, Xin Wang, Chaoyu Guan, Yateng Tang 等WWW 2022 · 被引用 50 次
- Dynamic Heterogeneous Graph Attention Neural Architecture SearchZeyang Zhang, Ziwei Zhang, Xin Wang, Yijian Qin 等AAAI 2023 · 被引用 44 次
- Adaptive Disentangled Transformer for Sequential RecommendationYipeng Zhang, Xin Wang, Hong Chen, Wenwu ZhuKDD 2023 · 被引用 32 次
- Large-Scale Graph Neural Architecture SearchChaoyu Guan, Xin Wang, Hong Chen, Ziwei Zhang 等ICML 2022 · 被引用 27 次
它引用的顶会 Paper13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
- Hyper-SAGNN: a self-attention based graph neural network for hypergraphsRuochi Zhang, Yuesong Zou, Jian MaICLR 2020 · 被引用 228 次
- Bridging the Gap between Sample-based and One-shot Neural Architecture Search with BONASHan Shi, Renjie Pi, Hang Xu, Zhenguo Li 等NeurIPS 2020 · 被引用 148 次
- AtomNAS: Fine-Grained End-to-End Neural Architecture SearchJieru Mei, Yingwei Li, Xiaochen Lian, Xiaojie Jin 等ICLR 2020 · 被引用 110 次
相关 Paper
- TextNAS: A Neural Architecture Search Space Tailored for Text RepresentationYujing Wang, Yaming Yang, Yiren Chen, Jing Bai 等AAAI 2020 · 被引用 66 次
- PSP: Progressive Space Pruning for Efficient Graph Neural Architecture SearchGuanghui Zhu, Wenjie Wang, Zhuoer Xu, Feng Cheng 等ICDE 2022 · 被引用 5 次
- Auto Learning AttentionBenteng Ma, Jing Zhang, Yong Xia, Dacheng TaoNeurIPS 2020 · 被引用 28 次
- K-armed Bandit based Multi-Modal Network Architecture Search for Visual Question AnsweringYiyi Zhou, Rongrong Ji, Xiaoshuai Sun, Gen Luo 等ACM MM 2020 · 被引用 25 次
- AutoGT: Automated Graph Transformer Architecture SearchZizhao Zhang, Xin Wang, Chaoyu Guan, Ziwei Zhang 等ICLR 2023
